Speech data · Language and dialect coverage

Which languages, which dialect regions under them. And what each speaker record carries.

150+ languages delivered, all five Nordic national languages by dialect region, every European language including the low-resource ones, MENA and selected global languages.

Pool depth per language and cell is confirmed in writing before you sign.

This page Nordic Active pool European Delivered MENA and global Delivered Pairs and new Scoped on request Consent · GDPR Article 7, per record

Coverage, by tier five tiers × six properties

A language label says little. What each tier carries is the map.

Every cell is a property of the corpus. A filled cell holds in every corpus of that tier. An outlined cell is confirmed per corpus at scoping. A dashed cell is outside the tier.

A tier marked Delivered means prior collection in that language. A region, cohort or setting within it is still confirmed per engagement.

Nordic depth five languages · the regions under them

Nordic is where the general models fail hardest. The regions we recruit by, under each language.

Regions are the standard dialect groups. Each becomes a quota cell when your corpus needs it, recruited in region and reviewed by a native linguist for that region.

Language 01 · nb-NO · nn-NO

Norwegian

Two written norms, five spoken regions. The norm is chosen per corpus; the region is a quota cell.

  • Østnorsk · Oslo and the east
  • Vestlandsk · Bergen, Stavanger
  • Trøndersk · Trondheim
  • Nordnorsk · Bodø to Finnmark
  • Sørlandsk · Kristiansand

Language 02 · sv-SE · sv-FI

Swedish

Six regional groups, Finland-Swedish scoped as its own cell.

  • Sveamål · Stockholm, Mälardalen
  • Götamål · Gothenburg and the west
  • Sydsvenska · Skåne
  • Norrländska
  • Gotländska
  • Finland-Swedish

Language 03 · da-DK

Danish

Three regional groups and the Copenhagen standard.

  • Jysk · Jutland
  • Ømål · Zealand and Funen
  • Bornholmsk
  • Copenhagen standard

Language 04 · fi-FI

Finnish

Western and eastern dialect groups and the Helsinki standard.

  • Western dialect group
  • Eastern dialect group
  • Helsinki standard

Language 05 · is-IS

Icelandic

One national variety. Age and setting are the axes that matter.

  • One national variety
  • Quota by age band and setting

Why the region matters. Whisper large-v3, word error rate, published.

  • 6.8% Norwegian Bokmål read speech, NST
  • 30.0% Norwegian Nynorsk Common Voice
  • 9.5% Swedish Common Voice
  • 11.3% Swedish NST

Trained on Norwegian and Swedish speech, NB-Whisper and KB-Whisper bring those to 2.2%, 12.6%, 4.1% and 5.2%. Same language, one written norm or one recording setting apart.

Kummervold et al., Interspeech 2024, arXiv 2402.01917 · KB-Whisper, National Library of Sweden, Interspeech 2025, arXiv 2505.17538 YPAI's own per-dialect condition set, Bergen to Tromsø, Skåne and Jutland →

The speaker record seven fields · every recording

Coverage is only real if it survives into the metadata.

Every recording carries these fields. They are what the Croissant 1.1 description and the data card are generated from.

The region on the record is the one a native linguist verified.
Language tag
BCP 47, with region subtag (nb-NO, nn-NO, sv-FI)
Dialect region
Self-reported at intake, verified by a native linguist for the region
First language
L1 or L2 speaker of the corpus language, with L1 named
Age band and gender
As the quota table defines them
Setting and device
Room, vehicle, ward or street; microphone class and sample rate
Consent
Consent id with purpose, scope and timestamp, GDPR Article 7
Integrity
SHA-256 per file, guideline version, capture timestamp

Review per region. Transcripts are checked by native linguists for the dialect region. Agreement is measured per dialect group, and disagreements are adjudicated by a dialect-specialist reviewer before the corpus is used for training or evaluation.

Pairs and the long tail scoped as a unit

A pair is the unit for code-switching. A pool snapshot is the first deliverable for a low-resource language.

Code-switching is collected as a pair, and the pair is what we scope: Norwegian and English in a workplace, Arabic and French, Swedish and Finnish. Speakers are recruited for the pair, prompts are built to provoke the switch, and transcripts carry a language tag per token so the switch points are visible to the model.

For a low-resource language the first deliverable is the pool snapshot: how many qualified speakers per region we can contact today, how many we can recruit within the plan. That number goes into the statement of work, and the corpus is sized to it rather than to a catalogue figure.

From the list to the quota table send · snapshot · quota · pilot · production

You send the language list. The pool snapshot comes back per cell, in writing.

Five stations between your list and a production corpus. The snapshot is the one that decides the rest, and it comes before signature.

  1. You send the list

    Languages, dialect regions, speaker cohorts, volumes, and the use.

  2. Pool snapshot

    Speakers per language and cell we can contact today, dated, in writing, before signature.

  3. Quota table

    Every cell with its target and the 5-point tolerance, BCP 47 tags and regions in the specification.

  4. Pilot

    Within 10 business days on the agreed cells, with the per-region QA report.

  5. Production

    Weekly deliveries, per-cell quota compliance, versioned manifest, Croissant 1.1 description and data card.

Three things stay explicit.

  • The 150+ figure is delivery history and network reach. Capacity for your corpus is the pool snapshot, answered per cell and in writing.
  • A cell that comes up short is named in the snapshot, with the recruitment plan to close it.
  • Languages outside the map are scoped from the pool check, and the answer can be a smaller corpus or a longer plan than requested.

Questions

Languages, dialects and the list

Do you publish a full language list?

The map above gives the tiers, and the Nordic frames name the regions. The per-language answer is the pool snapshot: speaker counts per cell, dated and in writing, during scoping. That number moves month to month, so it is answered fresh each time.

Can you cover Norwegian dialects beyond Oslo?

Yes. Vestlandsk, Trøndersk, Nordnorsk and Sørlandsk are quota cells, each recruited in region and each reviewed by a native linguist for that region. Error rates come back per region, so a corpus that is strong on Oslo and weak on Tromsø shows as exactly that.

How do you tell a speaker's dialect region?

Two ways, both recorded. The speaker reports region and first language at intake, and a native linguist for the region verifies it on the first recordings. Both fields stay on the record, so a buyer can see where they agree.

What does "error rate per dialect group" mean in practice?

Evaluation results are reported for each dialect group in the corpus separately, never as one blended average. The published Norwegian figures above show why: a single number would hide a 6.8% to 30% spread inside one language.

What if our language is not on the map?

It is scoped from the pool check. We tell you how many qualified speakers per region we can contact now and how many we can recruit within the plan, and the corpus is sized to that number.

How is code-switching collected?

As a language pair. Bilingual speakers are recruited for the pair, prompts are designed to provoke natural switching, and the transcript carries a language tag per token.

Bring the language list. The snapshot comes back per cell.

Languages, dialect regions, cohorts and volumes are enough to start. The quota table, the pilot and the per-region QA report follow from that.