Has the corpus ever been in a cabin?
A highway cabin runs at about 70 dB SPL: engine drone, HVAC, tyres, a passenger talking over the wake word. Booth-recorded corpora flatter the model in evaluation and ship false rejects to the driver on day one.
AI data, evaluation and implementation under one accountable delivery model.
Contact us Become a ContributorAutomotive voice and audio data
Wake-word, command, conversational and driver-monitoring audio, recorded by native speakers against real engine, HVAC and road noise, and annotated on the same engagement.
DPA per engagement · EEA processing by default
Signal What ships: accepted, labelled speech.
Noise What the cabin adds: engine, HVAC, road, overlap.
Native review A named human decision on every batch.
CH 01Why in-cabin voice fails
Three questions decide whether a voice programme survives contact with the road.
A highway cabin runs at about 70 dB SPL: engine drone, HVAC, tyres, a passenger talking over the wake word. Booth-recorded corpora flatter the model in evaluation and ship false rejects to the driver on day one.
Nordic dialects, accent drift and code-switched cabin conversations are the long tail that breaks the corpus built for the first five. We record them in-region with native speakers, not with synthetic accent transfer.
In-cabin voice classifies as high-risk under the EU AI Act, and Article 10 comes down to lineage, representativeness and bias examination. We ship the evidence pack at delivery, not on subpoena.
CH 02One session, end to end
Session 0412, from a live programme shape: captured in motion, annotated to your taxonomy, reviewed by a native speaker, accepted into a versioned batch.
Accepted material moves into versioned batches, manifests and secure delivery.
Drivers and passengers in real cabins. Engine load, road surface, HVAC and passenger overlap are part of the recording, not an afterthought.
Wake-word ground truth, intent and entities, speaker and seat-position metadata, labelled against the programme's schema.
Native human review is platform core. Failed recordings return for re-recording, rework or replacement according to the project rules.
CH 03The capture protocol
We scope the people, vehicles and conditions that decide whether voice works outside the lab. Every session carries the metadata to prove where it was made.
What the cabin adds
What the corpus must carry
Protocol, participant profile and review criteria are locked before production begins.
CH 04Language coverage
Native speakers, recruited where your buyers drive. The exact language list is set with the programme at scoping.
Nordic
Swedish Norwegian Finnish Danish Icelandic
Western Europe
French German Spanish Italian Dutch Portuguese
Eastern Europe
Polish Romanian Czech Hungarian Bulgarian
Asia Pacific
Mandarin Japanese Korean Thai
Regional varieties are in scope where the market needs them: Swiss German and regional German varieties, French variants for Belgium, Switzerland and Quebec, and Scottish, Welsh and Irish English accents.
Evaluation engagements, data-collection programs and active annotation work across OEM and Tier-1 suppliers.
CH 05Programme shapes
Market, speaker, cabin, rights and acceptance requirements turn scope into a programme specification before anything is recorded.
Positive and negative wake-word examples, in-domain commands and market-specific pronunciation, with false-accept and false-reject sets built to the product spec.
Ground-truth set · per market
Intent taxonomies aligned to your NLU schema, slot and entity extraction, multi-turn interaction and code-switched cabin conversations.
Annotated turn
start route guidance to the office
Native-speaker and professional-voice sourcing, style and prosody requirements, with project-specific model-training and deployment rights.
Voice specification
Speaker and seat-position metadata, overlapping cabin events and audio-event annotation, paired with gaze ground truth where the DMS spec requires both.
Audio-event classes
ASR output review, word and intent error analysis, and gap collection for underperforming markets or conditions.
Error analysis
CH 06How delivery runs
One controlled operating record. Reviewers and operations teams see the same task status and version history.
The forecast uses the unit that matters to the programme: accepted speaker-hours, recordings, sessions or assets.
CH 08Delivery record
Native-speaker recording across 50+ target-market languages in real cabin noise, captured by a vetted contributor network under GDPR-tracked consent. Wake-word ground truth and intent labels delivered with the audio. Output is a reusable corpus for in-cabin assistant fine-tuning, not a one-shot dataset.
Per-OEM branded wake word with false-accept and false-reject sets, an intent taxonomy aligned to the buyer's NLU schema, and command classification across target-market languages. Ground truth versioned alongside your on-device model so retrain cycles do not lose lineage.
Drowsiness, distraction and emergency audio classes captured and labelled with severity tiers, paired with gaze and eyelid ground truth where the DMS spec requires both modalities. Project taxonomies can follow Euro NCAP 2026 in-cabin protocols and defined edge-case coverage.
CH 09The full automotive surface
The full automotive data surface in one engagement: in-cabin voice corpora, wake-word and DMS audio, perception annotation, real-world capture. EEA jurisdiction across all of it.
CH 10Scope a programme
Share the actual modalities, target markets, participant profile, capture environment, output schema and acceptance criteria. The pilot and production path are defined around those requirements.
Build
Cabin scenarios, speaker profiles, scripts and audio requirements become a capture and annotation plan, delivered as an accepted corpus.
Remediate
Model outputs and failure cases isolate gaps by market, speaker profile or cabin condition, then drive targeted evaluation or collection.
Produce
A defined brief, rights model and acceptance criteria become qualification, calibration, production, QA, rework and accepted batches.
A named EU-resident project lead replies with feasibility, language coverage and a first read on Article 10 risk classification.
Target languages, capture profile, wake-word spec and the Article 10 evidence-pack manifest, as an indicative scope.
The actual modality, markets, participant profile, cabin conditions, output schema and acceptance criteria agreed for the programme.
Processing locations, sub-processors, delivery plan and production acceptance gates agreed before scale-up.
Start with a scoped pilot Modality, markets, volumes, QA and commercial terms, defined from your brief.
Norwegian Aksjeselskap. EEA-resident operations. GDPR Article 7 consent on every contributor. EU AI Act Article 10 evidence pack at delivery.
CH 11Questions
The questions an automotive voice or data lead asks before scoping a programme.