SPEECH DATA FOR ASR, TTS AND VOICE AI. BUILT IN THE EEA

Same test, Norwegian training. 36% lower WER out of domain.

Speech models fail where their training data never went. Dialect, spontaneous speech, code-switching.

NB-Whisper large 6.6% vs OpenAI Whisper large-v3 10.4% · FLEURS Bokmål · Kummervold et al., Interspeech 2024

Condition

Move across the frame to scrub the take

THE CONDITION SPACE Device · Setting · Speakers · Speech type

Four axes decide whether your speech data survives deployment

You name the range your system has to hold. The corpus is built to span it, and every recording lands with its coverage recorded.

AXIS 01

Device

The microphone in the user's hand, not the one in the studio. A model trained on one tier degrades on the next.

  • Smartphone tier
  • USB headset
  • In-vehicle microphone
  • Lavalier and body-worn
  • Far-field array
  • Studio condenser

AXIS 02

Setting

The room, the vehicle, the street. Noise, reverberation and competing speech are specified, not hoped for.

  • Studio
  • Home
  • Contact centre
  • Automotive cabin
  • Clinic and ward
  • Street and outdoor

AXIS 03

Speakers

Recruited by dialect region, age and background, so the long tail is represented rather than averaged away.

  • Dialect region
  • Age range
  • Gender identity
  • Second-language accent
  • 150+ languages, Nordic and low-resource depth

AXIS 04

Speech type

Targeted capture of the speech a benchmark removes and production does not.

  • Read and prompted
  • Spontaneous
  • Multi-speaker and overlapping
  • Code-switched mid-sentence
  • Wake word and command
  • Emotional and expressive

Coverage is agreed before recording begins and delivered as part of the record. Where a required range cannot be covered, YPAI says so during scoping rather than after delivery.

COVERAGE, RECORDED Device × Setting · 36 conditions

A benchmark covers one corner. A corpus covers the grid

Every cell is a named condition, device by setting. Public benchmarks cluster in the quiet corner: read speech, one microphone, one dialect. Your corpus is specified against the cells your users will actually speak in, and ships the proof in the record.

RUN THE THESIS The quiet room · The world

Hear it on your own voice

Record three seconds in your quiet room. The same take is transcribed clean, and again with the world mixed in at a chosen level. The words that survive are the difference training data has to cover.

The quiet room reference

startrouteguidancetotheharbourterminal

The world Street · traffic

stoprootguidancetoharbourtunnel

Runs in this tab · nothing leaves your device Specimen shown until you run it · illustrative, no customer data

WHICH SITUATION ARE YOU IN Collection · Voice · Labels · Evaluation · Licensing · Specification

Choose the work that matches the gap

Six paths, one operation. Each names what it produces and where its depth lives, and no path ships without the same record behind it.

  1. The speech you need does not exist yet

    YPAI designs the collection, recruits and qualifies the speakers, and runs the recording operation. Read, prompted, spontaneous, conversational, command or wake-word speech, remote or in studio, delivered against agreed speakers, conditions and checks.

    Coverage and recruitment
  2. One voice matters more than population coverage

    Specialist voice recording defines casting, pronunciation, performance and recording conditions. Recorded voice files with model-use rights stated in writing; voice-cloning, buyout, exclusivity and resale rights need an explicit written grant.

    Voice recording and rights
  3. You have audio, but the model cannot use it

    Labelled corpora pair source audio with an agreed schema and review standard: transcription, segmentation, diarisation, alignment and linguistic annotation, checked by native reviewers with the QA trail kept.

    Transcription, diarisation and labels
  4. Locate the failure before adding data

    Evaluation packages pair a set, rubric and reviewer process, then separate results by language, dialect, device, setting or speaker group. The output can include an error taxonomy and a regression set your team can run again.

    Evaluation sets and regression
  5. Could an existing dataset fit?

    YPAI reviews an existing or partner-sourced dataset against intended use, speech type, acoustic conditions, format, provenance and permitted rights. Item-specific licence, gap-fill scope, or both.

    Dataset review and licensing
  6. Your receiving team needs a definition of correct delivery

    A technical specification for the receiving system: formats, channel and track structure, transcript schema, metadata, validation and the conditions for acceptance.

    Technical specification

Specialist collection · 15 Explore collection

  1. Nordic and low-resource languages
  2. Dialect and accent corpora
  3. Conversational and multi-speaker audio
  4. Full-duplex dialogue for voice agents
  5. Code-switching speech
  6. In-car and automotive voice
  7. Contact-centre dialogue
  8. Clinical and healthcare speech
  9. Wake words and voice commands
  10. Far-field and noisy environments
  11. TTS voice production
  12. Emotional and expressive speech
  13. Studio, field and in-person capture
  14. Parallel corpora and MTPE, 300+ language pairs
  15. Ready-made speech datasets

A speech type or setting not listed here? YPAI designs the capture protocol around the deployment; describe the system and the condition it runs under.

Scope a corpus

HOW A CORPUS IS BUILT Coverage specification · Capture · Quality gates · Delivery

Every recording arrives with its record

One operation, four stations. Each station stamps the record that travels with the audio, so every delivery arrives ready to inspect.

One recording. Lavalier and body-worn · Street and outdoor · Spontaneous

  1. Coverage specification

    The four axes are turned into quotas: which dialects, which settings, which devices, and how much of the edge. Pool depth per language and accent is confirmed before signature.

    • Quotas Held within 5 percentage points per cell
    • Acceptance Criteria written before recording starts
  2. Capture

    Identity-verified speakers record under the specified conditions, consent per recording.

    • Studio 48 kHz / 24-bit, SNR 40 dB or better
    • Field 16 kHz or above, SNR and environment class per file
  3. Quality gates

    Automated technical QC on every file, then native-reviewer transcription checks against the agreed sampling plan.

    • Floor 98% first-delivery acceptance, re-collected below it
    • Evaluation WER per dialect group, never one blended number
  4. Delivery

    Audio, transcripts and metadata in the formats and location you named, with a signed manifest.

    • Pilot Output within 10 business days of kick-off
    • Destination S3, GCS, Azure, SFTP or on-premise

Formats, channel and track structure, transcript schema, validation and acceptance are described in full. Read the technical specification →

WHO SPEAKS Recruited · Identity-verified · Qualified · Calibrated · Reviewed

Qualified speakers, not a crowd

YPAI is not a crowdsourcing platform: every speaker is recruited, identity-verified and qualified for the specific corpus they record in, by dialect region where the corpus needs it.

  1. Recruited

    Sourced for the corpus's languages, dialect regions, demographics and settings

  2. Identity-verified

    Verified identity on file, consent per recording

  3. Qualified

    Passes the corpus's own screening recordings before production

  4. Calibrated

    Transcript agreement tracked per dialect group during production

  5. Reviewed

    Independent native-linguist review on sampled work

Corpora draw on a network of 210,000+ registered contributors across 50+ countries; each mobilises only the qualified slice it needs.

CONSENT, RIGHTS AND RESIDENCY Legal basis · Consent · Rights · Provenance · Residency · Agreements

The record your legal and security review will ask for

A recording can meet its technical requirements while its permission does not cover the next intended use. Rights, evidence and delivery stay connected, produced by the operation rather than assembled afterwards.

Where a person is in the recording, consent travels with the file.

The review asks · the record answers

  1. 01 What is the legal basis for each recording?

    Legal basis Consent, per person, per purpose and per session, documented per recording

  2. 02 Can you show consent, per speaker, and what happens on withdrawal?

    Consent GDPR Article 7, collected on the YPAI platform; withdrawals quarantined within 72 hours, reported within one business day, erased within 30 days

  3. 03 Which uses does the release cover?

    Rights Evaluation, model-training and commercial-deployment permissions defined separately; voice-cloning and resale only by explicit written grant

  4. 04 Where did each recording come from?

    Provenance Consent record, capture metadata and SHA-256 hash per file; Croissant 1.1 description and a data card in EU AI Act Article 10 order per dataset

  5. 05 Where does the audio reside, and under whose jurisdiction?

    Residency EEA by default, Norwegian jurisdiction; EEA-only sourcing and processing on request

  6. 06 What governs the engagement on paper?

    Agreements DPA signed, SCCs available for cross-border transfer

Every decision stays bound to the frame.

YPAI Data Collection & Assurance Platform The record of one frame

Every delivery carries the records the engagement requires and an integrity check on what is delivered.

WHERE THE SPEECH HAPPENS Healthcare · Automotive · AI and model companies · Public sector · Industrial and energy · Education

The vocabulary, the room and the rules change with the industry

A clinic, a cabin and a contact centre do not sound alike, and the consent, residency and acceptance rules differ as much as the acoustics. Six industries carry their own speech pages; every other one starts at the intake below.

All industry solutions →
  1. Healthcare

    Clinical speech for ambient documentation and dictation in Nordic languages, recorded under consent frameworks a hospital review can read.

    EEA-resident, consent per recording

    Healthcare
  2. Automotive

    In-cabin voice across dialects, road noise and microphone positions, for wake words, commands and multilingual driver interaction.

    Cabin, city and motorway conditions

    Automotive
  3. AI and model companies

    Training and evaluation corpora for ASR, TTS and speech-to-speech models, with dialect-balanced sets and results separated per condition.

    WER per dialect group, never blended

    AI and model companies
  4. Public sector

    Citizen-facing speech in Bokmål, Nynorsk and the dialects between them, for service lines, case dictation and accessibility, under Norwegian jurisdiction.

    EEA-resident, Norwegian jurisdiction

    Public sector
  5. Industrial and energy

    Voice on the floor and in the field, with hearing protection, machinery noise, radio and far-field capture for inspection and maintenance workflows.

    Far-field and noisy-environment capture

    Industrial and energy
  6. Education

    Learner speech across ages and accents, pronunciation and reading data for language learning and accessibility products.

    Nordic and European language breadth

    Education

PILOT TO PRODUCTION Collection · Annotation · Validation · Human and model evaluation · Accepted

One pilot against your requirement.

A pilot is scoped to your specification, quality thresholds and acceptance criteria, and reviewed against them before anything scales. Scope and commercial terms are agreed before it starts. Production is a separate decision, taken after the pilot review.

The pilot fixes

  • Scope of work
  • Technical requirements
  • Acceptance criteria
  • Data protection and rights
  • Commercial structure
  • Remediation and change control
  1. scope your specification and acceptance criteria
  2. terms scope and commercial terms agreed before it starts
  3. review against the agreed criteria, in a pilot workspace
  4. production a separate decision, taken after the review

Production: a separate decision, taken after the pilot review

Scope a pilot

ENGAGEMENT PROCESS

From the first brief to a production corpus

  1. 01 Describe the system What it has to do, where it fails, and how the resulting data will be used.
  2. 02 Feasibility read Languages, dialects, settings and volume checked against pool depth before anything is promised.
  3. 03 Pilot delivery Representative output within 10 business days to validate quality gates, formats and the record.
  4. 04 Production corpus Weekly deliveries with per-dialect QA, versioning and the signed manifest.

GDPR Article 7 · EU AI Act Article 10 · DPA included

The brief
Speech type (optional)

Include: languages and dialects, settings and devices, speech type, volume estimate, and any rights or residency constraints.

GOVERNANCE

Consent, rights and audit readiness

What governance artefacts can be delivered with a speech dataset?

Consent records per speaker, capture metadata and a SHA-256 hash per file, demographic and dialect breakdowns, QA audit trails, a Croissant 1.1 description and a data card written in the order EU AI Act Article 10 asks for, and a signed data processing agreement.

Can the same recordings be used for training, evaluation and a commercial product?

Only if the release says so. YPAI defines evaluation, model-training and commercial-deployment permissions separately, with territory, term, withdrawal handling, deletion and retention. If the recordings will feed models you have not named yet, that use is written in rather than assumed later.

What happens when a speaker withdraws consent?

The recordings are quarantined in YPAI systems within 72 hours, the withdrawal is reported to you within one business day with a ledger entry, and erasure completes within 30 days.

Where is the audio processed and stored?

Primary operations are in Norway on EEA infrastructure. Projects can be structured for EEA-only participant sourcing and EEA-resident processing; US or on-premise delivery is arranged where a deployment requires it.

How is transcription quality measured on dialectal speech?

Transcripts are checked by native linguists for the dialect region, agreement is measured per dialect group rather than as one blended average, and disagreements are adjudicated by a dialect-specialist reviewer before the corpus is used for training or evaluation.

Bring the system, the corpus or the performance gap.

If YPAI is not the right fit, we will say so directly.