EhkiLicensing inquiries →
For AI labs, API providers & voice AI startups

Voice data for the accents and languages your benchmarks aren't catching.

Foundation models are only as good as the data they train on, and the accents your models still get wrong are, by definition, not sitting in the open internet for your crawlers to find. Ehki collects English speech directly from native speakers of underrepresented language backgrounds, with explicit consent to train and commercially license it. Currently spanning speakers with Yoruba, Igbo, Hausa, and Nigerian Pidgin backgrounds, whose English carries the accent and code-switching patterns of those languages, with more backgrounds and regions on the roadmap.

Why this dataset, specifically

Not clean scripted reads bundled as a broad regional dataset

This is natural, spontaneous English speech from Yoruba, Igbo, Hausa, and Nigerian Pidgin speakers, including real code-switching into those languages mid-sentence: the messy, real-world speech that clean benchmark read-speech doesn't capture, and that production systems actually fail on.

One dataset, growing by speaker background and region

Every recording is tagged with the speaker's native language background and metadata (region, age range, fluency), so you can license the whole dataset or just the accent groups you need. New speaker backgrounds and regions are added over time, not bundled once and left static.

An ongoing relationship, not a one-time drop

Available under non-exclusive or exclusive licensing agreements, with quarterly data refreshes included, so your fine-tuning set doesn't go stale the way a static purchase does.

Provenance documentation isn't an afterthought

Every speaker explicitly consented to commercial AI training and redistribution use, with a timestamped, versioned consent record and per-recording provenance you can audit, the kind of documentation buyers now require before a deal closes, not something bolted on after. For our Nigerian speaker programs, that consent process is aligned with the Nigeria Data Protection Act (2023).

Every recording is manually reviewed before it ships

Submissions are checked for authenticity (natural, unfaked accents), audio quality, and consent completeness before being marked approved. Recordings that fail review are excluded from any delivered dataset, not shipped with a quality flag for you to filter out yourself.

Dataset snapshot

We're actively collecting and growing this dataset, refreshed quarterly. Full composition stats (hours, unique speakers, regional breakdown, code-switching coverage) are shared directly with buyers evaluating the data, so get in touch below.

Consent & compliance, in plain language

Every contributor explicitly agreed, before recording anything, that their voice may be used to train AI models and licensed commercially to third parties, including AI companies abroad. This was disclosed clearly at collection time, not buried in fine print, and each consent record is versioned and timestamped so we can show exactly what a speaker agreed to. Speaker names and directly identifying details are never included in any dataset delivered to a buyer; only anonymized metadata (region, age range, gender, language background) accompanies each recording.

Data format & delivery

Delivered as a versioned export in a private cloud storage bucket (S3-compatible), scoped to your licensing agreement, not a shared folder link. Each export includes:

Audio

16-bit PCM WAV, transcoded to your target sample rate (16kHz for ASR, 44.1kHz for TTS/voice work), one file per recording, named by an anonymized speaker and recording ID.

Metadata

A JSONL manifest with one row per recording (speaker ID, duration, transcript, recording environment, code-switch tags and languages, review status), plus a separate speaker table (anonymized ID, region, age range, gender, native language background, self-rated fluency). Parquet available on request. Scripted recordings ship with ground-truth transcripts, since the script text is known in advance; spontaneous recordings are transcribed separately and flagged as such, so you always know which is which.

Documentation

A datasheet covering motivation, composition, collection process, consent basis, and intended/excluded uses, in the format ML teams already expect from a serious data vendor, plus the license terms for that specific delivery.

A small sample export follows this same structure, so evaluating fit doesn't require a call first, just a look at the actual files and manifest.

Licensing options

This is a licensing relationship, not a one-time data drop. Pricing depends on scope, exclusivity, and how you plan to use the data.

Non-exclusive

License one or more speaker groups (by language background) with quarterly refreshes. Available to multiple buyers.

Exclusive

Exclusive rights to a speaker group or region, scoped by application or geography if you don't need full exclusivity.

Custom collection

Need a specific accent, region, or domain vocabulary we don't collect yet? We can scope a dedicated collection program for your use case.

Get in touch and we'll walk through pricing and terms based on your use case: licensing@ehki.ai.