Dialect is not the main story
Whisper large-v3 scores 6.8% Word Error Rate (WER) on standard Norwegian Bokmål read speech on the NST test set, a benchmark result that looks production-ready on paper. Give the same model Nynorsk speech from the Common Voice test set and WER climbs to 30% (Kummervold et al., Interspeech 2024). That is not a rounding error. That is nearly one in three words wrong, in the same language, from the same model.
Those are two different test sets, so the cleaner comparison comes from the National Library’s own dialect-tagged evaluation (Solberg et al., January 2024): 10 hours of spontaneous NRK radio and TV speech, 7,261 segments, 409 speaker instances representing 360 unique people, every one tagged with dialect region. On that audio, OpenAI’s Whisper large scores 27% to 32% WER across all five dialect regions with a Bokmål target, 30.2% on the full set, and 53.5% with a Nynorsk target. The larger measured gaps are read versus spontaneous speech, Bokmål versus Nynorsk target, and single-speaker versus overlapping speech.
Two cautions before that number does any work. The National Library cautions that WER is an imperfect measure for Whisper because the model does not always transcribe verbatim. A high WER can therefore overstate the loss in readability or meaning. The number remains operationally important where exact transcription is required, and the report also identifies genuine omissions, hallucinations and language errors. And the Interspeech result uses Whisper large-v3, while the NRK report labels its checkpoint only as openai-whisper-large and does not establish that it is large-v3. The comparison therefore demonstrates a model-family and test-distribution gap, not a controlled same-checkpoint experiment.
The published results strongly implicate training-data and evaluation-distribution mismatch. Targeted Norwegian training substantially reduces the gap, but the evidence does not show that model architecture is irrelevant. The gap is also unlikely to be Whisper-specific: any general-purpose Automatic Speech Recognition (ASR) model trained mostly on read and broadcast standard-variety speech should show it.
The Training Data Problem Behind the Benchmark
OpenAI trained the original Whisper on 680,000 hours of web-scraped audio; large-v3 raised that to roughly 1 million hours of weakly labeled audio plus 4 million hours pseudo-labeled by large-v2. That scale sounds exhaustive until you examine the distribution. Web-scraped speech data skews heavily toward English, and within non-English languages, it skews toward broadcast-quality, standard-dialect recordings, the kind of Norwegian spoken on NRK national radio, not in a Trøndersk fishing cooperative or a Northern Norwegian municipal office.
The result is a model that has learned Norwegian as it appears on the internet, not as it is spoken by the roughly 5.3 million people who speak it natively in daily life. Regional dialects, code-switching patterns, and spontaneous conversational speech are systematically underrepresented. Scandinavian languages are a clear case of this pattern, and the same dynamic plausibly affects Finnish, Danish regional varieties, and Swedish dialects outside the Stockholm standard.
Why This Is a Production Problem Right Now
This matters beyond academic benchmarks. Automotive, fintech and telehealth teams deploying Norwegian voice interfaces face the same exposure the benchmarks describe: demo conditions that resemble read broadcast speech, and field conditions that do not, once real users speak spontaneously, in real acoustic environments, in their own dialect.
The regulatory clock makes this concrete. The Digital Omnibus on AI, in force since July 27, 2026, moved the EU AI Act (Regulation 2024/1689) high-risk deadlines: Annex III systems now apply from December 2, 2027, and AI embedded in regulated products from August 2, 2028. A vehicle voice system is not automatically high-risk. The Annex I route applies only where the AI system satisfies the relevant product and safety-component classification conditions. Where the system is classified as high-risk and uses training, validation or test data, Article 10 requires documented data-governance practices for those datasets. The Article 50 transparency duties applied on August 2, 2026 as planned. A dialect gap you cannot explain is exactly the kind of finding an audit surfaces.
The following sections walk through the published evidence, examine what the data distribution underneath it actually looks like, and provide a practical framework for building speech corpora that narrow the WER gap at the source.
What the Published Benchmarks Cover, and What They Cannot
The peer-reviewed evaluation of Whisper on Norwegian is the National Library of Norway’s NB-Whisper work (Kummervold et al., Interspeech 2024). It measures OpenAI’s Whisper variants against three public test sets: NST (studio-quality Bokmål read speech), Fleurs (Bokmål), and Common Voice (Nynorsk). As of September 2026, large-v3 remains the strongest open-weights Whisper release; the speed-optimized, pruned large-v3-turbo checkpoint (decoder cut from 32 to 4 layers) trades a small amount of accuracy for faster inference.
Two things stand out in that paper. First, it measures the gap at the written-standard level (Bokmål versus Nynorsk), on three test sets with three different recording conditions. Second, the authors say a realistic picture needs test sets with speakers from different dialects and dialect metadata, which those three sets do not carry. The Nordic Dialect Corpus documents 38 distinct pronunciations of the interrogative “who” alone.
That metadata exists elsewhere. The National Library’s January 2024 evaluation, Status for norsk talegjenkjenning, built the test set the paper asks for: 10 hours of NRK radio and TV, 7,261 segments, every speaker tagged with one of five dialect regions and one of fifteen fine-grained dialects, plus gender, recording conditions and overlapping speech. It ran OpenAI Whisper large, three NB-Whisper variants, the library’s wav2vec2 models, Google’s USM and Cloud Speech, and Microsoft Azure against it.
| Dialect region | Whisper large, Bokmål | NB-Whisper large verbatim, Bokmål | Whisper large, Nynorsk | NB-Whisper large verbatim, Nynorsk |
|---|---|---|---|---|
| Northern Norway | 27.3% | 10.1% | 46.5% | 20.4% |
| South-western | 29.7% | 11.8% | 50.9% | 17.3% |
| Trøndersk | 30.1% | 12.0% | 54.2% | 26.7% |
| Western | 30.8% | 11.8% | 51.0% | 18.3% |
| Eastern | 31.7% | 13.2% | 59.9% | 23.4% |
WER on the full test set, lower is better. Source: Solberg et al., January 2024, figures 1 and 4.
Three findings reframe the dialect story. With a Bokmål target, Whisper’s spread across regions is under five points. Eastern Norway has the highest full-set WER in this sample, but the report attributes that result partly to sample composition and harder segments, not to Eastern speech being intrinsically more difficult; with the 25% hardest segments removed, Trøndersk becomes the hardest region for most models. With a Nynorsk target the spread is 13 points, and Trøndelag and Eastern Norway are hardest (Oslo and Trøndelag in the fine-grained view), which the authors attribute to Nynorsk training data most likely coming from the west. The fine-grained view shows 17-point gaps for Whisper (Østfold 21%, Midtlandsk 38%) that the authors warn rest on too few speakers to rank individual dialects.
The report’s own summary: dialect has some effect on WER, and the large effects come from elsewhere. Overlapping speech takes Whisper from 28% to 45% (Bokmål) and from 51% to 76% (Nynorsk). Background noise adds four points. And the gap between studio read speech (6.8%, large-v3) and spontaneous broadcast speech (30%, checkpoint not specified) is the largest of all, with the caveat above.
So the published numbers do not describe a hidden dialect problem waiting to be measured. They measure a spontaneous-speech problem that hits every dialect, an orthography problem that hits Nynorsk speakers hardest, and a conversation problem that a read-speech corpus is unlikely to fix.
The spoken dialect groups a production Norwegian corpus must cover, using the National Library’s five regions:
- Eastern Norwegian (Oslo, Innlandet, Østfold, Agder), the closest match to written Bokmål and the largest share of the NRK test set
- South-western (Rogaland), the region the Nynorsk models handle best in the report
- Western (Bergen, Sogn og Fjordane, Sunnmøre), Nynorsk-near speech, with Bergen scoring worse than the rest of the region in the fine-grained view
- Trøndersk (Trøndelag), the hardest region for most models once the hardest segments are removed, and the hardest for Nynorsk transcription
- Northern Norwegian (Nordland, Troms, Finnmark), the region with the lowest Whisper WER in this sample
The same structure repeats across Scandinavia: Skåne Swedish and Jutlandic Danish sit far from the Stockholm and Copenhagen varieties that dominate broadcast training data, and are the obvious candidates to stratify by in Swedish and Danish projects. The corpus framework later in this article generalizes accordingly.
Why Spontaneous Speech Matters More Than Read Speech
Read speech and spontaneous conversational speech are not the same task. The two Norwegian evaluations above show the size of the difference for the Whisper large family: 6.8% on studio read speech (NST, large-v3) against 30.2% on spontaneous broadcast speech (NRK, checkpoint not specified), before any in-cabin acoustic factors are introduced.
For in-cabin voice, the compounding is worse: active road noise, HVAC fan noise, multi-speaker overlap, natural hesitations, self-corrections, and mid-command dialect switches. A driver beginning a navigation command in standard Norwegian and completing it in Trøndersk is a realistic test case, not an exotic one.
The public Norwegian corpora underneath the benchmarks sit at the easy end of this spectrum. NST (Nordisk Språkteknologi) is studio-condition read speech, and the NB-Whisper team had to assemble parliamentary proceedings and NRK broadcast subtitles on top of it to get anywhere near real speech diversity. Neither distribution reflects what in-cabin ASR systems encounter at 110 km/h on the E6.
If your ASR training data corpus is 80% read speech from capital-city speakers, your benchmark results will not predict production performance. They will predict performance on a task your production system never actually faces.
Audio Annotation Protocol for Dialectal Speech
Dialectal speech annotation introduces problems that generic transcription pipelines are not designed to handle. The first is orthographic ambiguity: Trøndersk and Northern Norwegian have no standardized written form. An annotator transcribing a Trøndersk speaker saying what sounds like “kæm ær du” faces a genuine decision, transcribe in normalized Bokmål (“hvem er du”), attempt a phonetic approximation, or use a dialect-aware orthographic convention. Each choice has downstream consequences for ASR training data quality.
A workable convention for Norwegian dialect projects, and the one YPAI proposes as the default in a SOW, uses normalized Bokmål as the reference tier with a secondary tier for dialectal forms that have no Bokmål equivalent. This matches the NST corpus convention and allows WER calculation against a stable reference. The trade-off is that it understates the model’s phonological confusion, a Bokmål-normalized reference will not capture whether the model failed on a phoneme or a lexical form.
Expect annotator agreement to drop on dialectal audio relative to standard speech, which is why disagreements need adjudication by a dialect-specialist annotator and why agreement must be measured per dialect group, never as a blended average. Using general-purpose Norwegian or Danish speakers as annotators without dialect screening produces reference transcriptions with systematic errors, errors that propagate directly into WER calculations and, if the corpus is used for fine-tuning, into the model itself.
Results: Where Whisper Breaks Down and Why
The published results are consistent. Whisper large-v3 performs well on standard Bokmål read speech and degrades sharply as the input moves away from its training distribution. The degradation accelerates as models shrink.
| Test set | Whisper large-v3 | Whisper medium | Whisper small | NB-Whisper large |
|---|---|---|---|---|
| NST (Bokmål, read) | 6.8% | 14.6% | 27.2% | 2.2% |
| Fleurs (Bokmål) | 10.4% | 15.5% | 29.6% | 6.6% |
| Common Voice (Nynorsk) | 30.0% | 60.2% | >100% | 12.6% |
WER, lower is better; above 100% is possible when a model inserts more words than the reference contains. Source: Kummervold et al., Interspeech 2024.
Three failure modes are worth testing for separately in a deployment-specific evaluation set. Neither source breaks its errors down this way, so treat these as hypotheses to measure, not published findings.
Vocabulary gaps. Dialectal lexical forms with no Bokmål equivalent and little representation in web-scraped training data are candidates for substitution or deletion. Stratify the test set by dialect so this can be measured per region.
Phonological mapping. Where a dialect’s sound system diverges from the standard variety, a model may map to the nearest standard form it knows. Trøndersk in Norway and Jutlandic in Denmark are the obvious features to stratify by; whether they trigger systematic substitutions is what the evaluation has to show.
Language confusion. The most operationally damaging of the three if it occurs, addressed below.
Language Confusion: When Whisper Thinks Norwegian Is Swedish
Whisper identifies the language from the first 30 seconds of audio. For closely related languages, Norwegian, Swedish, Danish, the acoustic and lexical overlap is substantial, and misidentification on short utterances is a commonly reported problem among Whisper users, not a rate measured in either Norwegian evaluation.
When language ID is wrong, the decoder is conditioned on the wrong language and errors compound beyond the acoustic gap. NB-Whisper, the fine-tuned Norwegian model released by the National Library of Norway (Nasjonalbiblioteket), substantially reduces this confusion by retraining on Norwegian-specific data, and in the library’s per-dialect results its verbatim variant has the smallest spread across regions of any system tested. What no fine-tune can add is coverage of the conditions the public sets do not contain: short commands, cabin noise, and your users’ dialect mix.
Forcing the language tag via Whisper’s --language no flag removes one failure path. Whether the acoustic error remains after forcing is something a deployment-specific evaluation set has to measure. Treat language forcing as a control to test, not a fix to assume.
The Automotive Edge Case: Dialect + Noise + Short Utterances
The hardest real-world combination is utterances of a few words, ambient road and HVAC noise, and dialectal phonology, all simultaneously.
A driver saying slå på varmen (turn on the heat) in Trøndersk dialect, with HVAC fan noise at highway speed, is a fundamentally different acoustic signal than the same phrase spoken in Standard Bokmål in a quiet recording studio. The phonological form is different. The signal-to-noise ratio is different. And short commands give the 30-second language-identification window very little audio to work with.
No cited public benchmark measures this combination, which is exactly the problem: the conditions your product ships into are the conditions the public test sets do not cover. Given that Whisper large already sits at 30% on clean spontaneous speech, that noise adds four points and overlapping talk adds seventeen in the National Library’s test, and that medium and small collapse to 60% and beyond on Nynorsk, shipping an in-cabin dialect deployment without your own evaluation corpus means shipping blind.
Prompting and language-tag forcing do not address this. It requires ASR training data that reflects the actual acoustic conditions and dialectal distribution of the deployment environment. Vehicle telemetry (speed, HVAC state, window position, occupancy) is a candidate conditioning signal worth testing alongside the audio. That kind of domain-specific context does not exist in general-purpose speech corpora, and read-speech fine-tuning does not supply it.
Closing the Gap: Building Dialect-Aware Speech Corpora
The benchmark results above are not an argument against Whisper. They are an argument for building the right training data before deploying it. A structured approach to dialect-aware corpus construction is the most direct lever the published evidence supports for narrowing the WER gap, but only if the process is designed around the actual deployment conditions, not general-purpose speech collection norms.
Here is a five-step framework for building ASR training data that reflects dialectal reality.
Step 1: Dialect mapping. Before recruiting a single speaker, inventory the specific dialect groups your product must support. Weight them by user population and commercial priority, not by linguistic convenience. A Norwegian automotive voice interface deployed nationally must treat Northern Norwegian dialects as first-class targets, not edge cases. Document which dialects are in scope, which are out of scope, and why. This decision determines your collection budget and annotation requirements downstream.
Step 2: Speaker recruitment. Recruit native dialect speakers, not standard-dialect speakers asked to “speak naturally.” The phonological differences between Standard Bokmål and Trøndersk are not stylistic; they are structural. Standard-dialect speakers cannot produce them reliably on demand. Within each dialect group, recruit across age cohorts, gender, and sociolect. A corpus built exclusively from 25–40 year-old urban speakers will underperform on elderly rural speakers, and that failure will surface in production.
Step 3: Recording environment realism. For automotive AI data, record in actual vehicles under real road conditions, not anechoic chambers or quiet offices. Capture HVAC noise at multiple fan speeds, road noise at highway and urban speeds, and window configurations. For telehealth applications, record with consumer-grade microphones in home environments with representative background noise profiles. The acoustic conditions in your corpus must match the acoustic conditions in your deployment environment. Any gap between the two is a gap in model performance.
Step 4: Annotation with dialect expertise. Assign annotators who are native to each dialect region. Establish transcription conventions before annotation begins, decisions about how to represent dialect-specific phonology, code-switching, and non-standard orthography must be made once and applied consistently. Measure inter-annotator agreement per dialect group separately. A corpus where annotators disagree on 15% of tokens in Northern Norwegian speech is not a 15% quality problem; it is a systematic bias that will propagate through fine-tuning.
Step 5: Iterative fine-tuning and evaluation. Fine-tune your target ASR model on the new corpus, then evaluate per-dialect WER separately, not as a blended headline number. An acceptable blended score can conceal severe failure on a dialect group that represents a material share of users. Identify remaining high-error dialect groups and feed them into the next collection cycle. This is not a one-time project; it is a pipeline.
How Much Dialect Data Do You Actually Need?
The NB-Whisper model, released by the National Library of Norway (Nasjonalbiblioteket), demonstrates what targeted corpus investment produces. The peer-reviewed paper reports 22,184 source hours in the first training stage and 6,078 hours after cleaning in the second, assembled from NST, parliamentary proceedings, NRK broadcast subtitles and audiobooks. The National Library’s report and the model card use different accounting (around 50,000 hours and 8 million 30-second samples respectively) and are not directly comparable to unique source duration. The model cuts Whisper large-v3’s WER from 30% to 12.6% on Common Voice Nynorsk and from 6.8% to 2.2% on NST read speech (Interspeech 2024). On the National Library’s spontaneous NRK test, the verbatim variant cuts Whisper large from 30.2% to 11.9% (Bokmål) and from 53.5% to 21.6% (Nynorsk), and its spread across dialect regions is the smallest of any system tested. That Bokmål result is a 60.6% lower measured WER, the strongest in the January 2024 report and the one Språkstatus 2025 cites. The standard NB-Whisper large, closer to Whisper’s edited transcription style, scores 18.7% on the same set.
One limitation applies to that comparison. NB-Whisper was trained extensively on Norwegian parliamentary, NRK, audiobook and read-speech material. The National Library notes that the model may have encountered some underlying NRK audio and that its broadcast training distribution likely benefits performance on this NRK-based evaluation; the authors judged a large overlap effect unlikely. WER also rewards the verbatim variant’s literal transcription style, which the report says can raise the measured error of edited-style systems whose output is still readable. The report attributes most of the gain to the Norwegian training data, while noting it has no controlled model pair in which training data is the only difference. The comparison remains highly relevant, but it is not a clean out-of-domain ablation of data alone.
YPAI planning assumption, not a published result: you do not need 22,000 hours to move your metrics. Targeted corpora in the tens to low hundreds of hours can produce consequential WER reductions when the data matches the deployment distribution. That match, not raw volume, is the variable we scope against.
What we do not scope: adding 500 hours of standard-dialect read speech. This approach may improve headline WER on clean benchmark sets while leaving dialect-specific error rates unchanged. The model learns more of what it already knows. Annotation quality compounds this dynamic; our working assumption is that 50 hours with consistent, dialect-aware transcription outperforms 200 hours with inconsistent annotation, and the first per-dialect evaluation is where that assumption gets tested.
The planning target YPAI uses for a production-grade dialect-aware corpus is 50-200 hours per dialect group, sourced from spontaneous speech in realistic acoustic conditions, with annotation handled by dialect-native contributors working from documented transcription conventions.
Compliance Requirements for Nordic Speech Data Collection
Speech data collected in EU and EEA jurisdictions is not generic data. An identifiable voice recording is personal data. Where specific technical processing is used to uniquely identify a speaker, which is what voiceprints, speaker verification and identity-based diarisation do, the resulting biometric data falls within Article 9. The controller then needs an Article 6 lawful basis and an applicable Article 9(2) condition, which may be explicit consent. Where consent is the basis, Article 7 requires it to be freely given, specific, informed and unambiguous, and each speaker must understand the purpose of the recording, how long it will be retained, whether it will train commercial AI systems, and how to withdraw. The appropriate basis, notice, rights process and retention controls depend on the project and the roles of the parties.
EU AI Act Article 10 adds a second layer. A vehicle voice system is not automatically a high-risk AI system under Regulation 2024/1689. The Annex I route applies only where the AI system satisfies the relevant product and safety-component classification conditions (obligations from August 2, 2028); Annex III use cases apply from December 2, 2027. Where the system is classified as high-risk and uses training, validation or test data, Article 10 requires documented data-governance practices for those datasets, covering data sourcing methodology, annotation processes, known limitations, and quality assurance procedures. This documentation must be maintained throughout the system lifecycle, not assembled retroactively before an audit.
The practical implication: every speaker in your speech corpus needs a documented lawful basis and rights process covering purpose, retention period and how to exercise their rights. Data provenance, the chain of custody from recording session through annotation through model training, must be auditable. A corpus collected without a valid lawful basis, the required Article 9 condition where applicable, and appropriate transparency and rights controls may be unlawful to process, regardless of its acoustic quality.
Building compliance into corpus design from the first recording session is materially less expensive than retrofitting it after the fact. It is also what enterprise buyers in European markets ask about first.
Build a Scandinavian Speech Corpus That Actually Works
Closing the WER gap on Norwegian dialects, Swedish regional speech, or Danish spontaneous conversation requires training data that was collected with intent: dialect-stratified speaker recruitment, documented lawful-basis and rights processes, and dialect-native review.
YPAI can design and operate dialect-stratified speech collection and annotation projects with speaker criteria, deployment-representative recording conditions, project-specific transcription conventions, dialect-native review, acceptance thresholds and versioned delivery defined in the SOW, with the data-governance documentation Article 10 asks for produced as part of delivery. Coverage spans 150+ languages, including all Nordic languages.
For the full corpus build process, see the guide to speech corpus collection for enterprise ASR; for how the engine choice interacts with corpus strategy, the ASR software comparison. And for every published dialect WER result across European languages, not just Scandinavia, see our quarterly European Dialect ASR Benchmark.
Explore the speech data family, audio and speech annotation, the managed data collection operation, or contact us to scope a Nordic speech data project.
Frequently Asked