Speech data · Language and dialect coverage
Which languages, which dialect regions under them. And what each speaker record carries.
150+ languages delivered, all five Nordic national languages by dialect region, every European language including the low-resource ones, MENA and selected global languages.
Pool depth per language and cell is confirmed in writing before you sign.
This page Nordic Active pool European Delivered MENA and global Delivered Pairs and new Scoped on request Consent · GDPR Article 7, per record
Coverage, by tier five tiers × six properties
A language label says little. What each tier carries is the map.
Every cell is a property of the corpus. A filled cell holds in every corpus of that tier. An outlined cell is confirmed per corpus at scoping. A dashed cell is outside the tier.
A tier marked Delivered means prior collection in that language. A region, cohort or setting within it is still confirmed per engagement.
Nordic depth five languages · the regions under them
Nordic is where the general models fail hardest. The regions we recruit by, under each language.
Regions are the standard dialect groups. Each becomes a quota cell when your corpus needs it, recruited in region and reviewed by a native linguist for that region.
Language 01 · nb-NO · nn-NO
Norwegian
Two written norms, five spoken regions. The norm is chosen per corpus; the region is a quota cell.
- Østnorsk · Oslo and the east
- Vestlandsk · Bergen, Stavanger
- Trøndersk · Trondheim
- Nordnorsk · Bodø to Finnmark
- Sørlandsk · Kristiansand
Language 02 · sv-SE · sv-FI
Swedish
Six regional groups, Finland-Swedish scoped as its own cell.
- Sveamål · Stockholm, Mälardalen
- Götamål · Gothenburg and the west
- Sydsvenska · Skåne
- Norrländska
- Gotländska
- Finland-Swedish
Language 03 · da-DK
Danish
Three regional groups and the Copenhagen standard.
- Jysk · Jutland
- Ømål · Zealand and Funen
- Bornholmsk
- Copenhagen standard
Language 04 · fi-FI
Finnish
Western and eastern dialect groups and the Helsinki standard.
- Western dialect group
- Eastern dialect group
- Helsinki standard
Language 05 · is-IS
Icelandic
One national variety. Age and setting are the axes that matter.
- One national variety
- Quota by age band and setting
Why the region matters. Whisper large-v3, word error rate, published.
- 6.8% Norwegian Bokmål read speech, NST
- 30.0% Norwegian Nynorsk Common Voice
- 9.5% Swedish Common Voice
- 11.3% Swedish NST
Trained on Norwegian and Swedish speech, NB-Whisper and KB-Whisper bring those to 2.2%, 12.6%, 4.1% and 5.2%. Same language, one written norm or one recording setting apart.
Kummervold et al., Interspeech 2024, arXiv 2402.01917 · KB-Whisper, National Library of Sweden, Interspeech 2025, arXiv 2505.17538 YPAI's own per-dialect condition set, Bergen to Tromsø, Skåne and Jutland →
The speaker record seven fields · every recording
Coverage is only real if it survives into the metadata.
Every recording carries these fields. They are what the Croissant 1.1 description and the data card are generated from.
- Language tag
- BCP 47, with region subtag (nb-NO, nn-NO, sv-FI)
- Dialect region
- Self-reported at intake, verified by a native linguist for the region
- First language
- L1 or L2 speaker of the corpus language, with L1 named
- Age band and gender
- As the quota table defines them
- Setting and device
- Room, vehicle, ward or street; microphone class and sample rate
- Consent
- Consent id with purpose, scope and timestamp, GDPR Article 7
- Integrity
- SHA-256 per file, guideline version, capture timestamp
Review per region. Transcripts are checked by native linguists for the dialect region. Agreement is measured per dialect group, and disagreements are adjudicated by a dialect-specialist reviewer before the corpus is used for training or evaluation.
Pairs and the long tail scoped as a unit
A pair is the unit for code-switching. A pool snapshot is the first deliverable for a low-resource language.
Code-switching is collected as a pair, and the pair is what we scope: Norwegian and English in a workplace, Arabic and French, Swedish and Finnish. Speakers are recruited for the pair, prompts are built to provoke the switch, and transcripts carry a language tag per token so the switch points are visible to the model.
For a low-resource language the first deliverable is the pool snapshot: how many qualified speakers per region we can contact today, how many we can recruit within the plan. That number goes into the statement of work, and the corpus is sized to it rather than to a catalogue figure.
From the list to the quota table send · snapshot · quota · pilot · production
You send the language list. The pool snapshot comes back per cell, in writing.
Five stations between your list and a production corpus. The snapshot is the one that decides the rest, and it comes before signature.
-
You send the list
Languages, dialect regions, speaker cohorts, volumes, and the use.
-
Pool snapshot
Speakers per language and cell we can contact today, dated, in writing, before signature.
-
Quota table
Every cell with its target and the 5-point tolerance, BCP 47 tags and regions in the specification.
-
Pilot
Within 10 business days on the agreed cells, with the per-region QA report.
-
Production
Weekly deliveries, per-cell quota compliance, versioned manifest, Croissant 1.1 description and data card.
Three things stay explicit.
- The 150+ figure is delivery history and network reach. Capacity for your corpus is the pool snapshot, answered per cell and in writing.
- A cell that comes up short is named in the snapshot, with the recruitment plan to close it.
- Languages outside the map are scoped from the pool check, and the answer can be a smaller corpus or a longer plan than requested.
Questions
Languages, dialects and the list
Do you publish a full language list?
The map above gives the tiers, and the Nordic frames name the regions. The per-language answer is the pool snapshot: speaker counts per cell, dated and in writing, during scoping. That number moves month to month, so it is answered fresh each time.
Can you cover Norwegian dialects beyond Oslo?
Yes. Vestlandsk, Trøndersk, Nordnorsk and Sørlandsk are quota cells, each recruited in region and each reviewed by a native linguist for that region. Error rates come back per region, so a corpus that is strong on Oslo and weak on Tromsø shows as exactly that.
How do you tell a speaker's dialect region?
Two ways, both recorded. The speaker reports region and first language at intake, and a native linguist for the region verifies it on the first recordings. Both fields stay on the record, so a buyer can see where they agree.
What does "error rate per dialect group" mean in practice?
Evaluation results are reported for each dialect group in the corpus separately, never as one blended average. The published Norwegian figures above show why: a single number would hide a 6.8% to 30% spread inside one language.
What if our language is not on the map?
It is scoped from the pool check. We tell you how many qualified speakers per region we can contact now and how many we can recruit within the plan, and the corpus is sized to that number.
How is code-switching collected?
As a language pair. Bilingual speakers are recruited for the pair, prompts are designed to provoke natural switching, and the transcript carries a language tag per token.
Bring the language list. The snapshot comes back per cell.
Languages, dialect regions, cohorts and volumes are enough to start. The quota table, the pilot and the per-region QA report follow from that.
Speech data overview Technical specifications Evaluation program Dialect benchmark Consent framework