An empty voice recording booth in an office at dusk, a phone on the ledge under one lamp.

SPEECH AND AUDIO INVENTORY

Speech that already exists. Read by the room it was recorded in.

Rights-cleared speech and audio YPAI owns or represents, qualified against your condition before it is offered. What is missing is collected as new recordings.

You describe the speech your system has to hold: the device, the setting, who speaks and how. YPAI reads its own and partner inventory against that condition, shows what can be licensed quickly under NDA, and records only what is missing. Every item arrives as a card your engineers and your counsel can read field by field.

the booth pictured, described
setting
a small booth in an ordinary office
device
a phone on the ledge, idle
speakers
one speaker, recruited by dialect region
speech type
spontaneous, prompted by a scenario
rights
consent on record, use set by the schedule

THE FOUR AXES

Four axes describe every item. Choose the ones your model runs under.

The same four axes the speech family collects against. Select a value on each and the card below fills for that condition. An axis you leave open stays open on the card, and YPAI collects new recordings for it.

axis Device

The microphone in the user's hand, not the one in the studio. An item is tagged with the tier it was captured on.

axis Setting

The room, the vehicle, the street. Noise, reverberation and competing speech are recorded on the item, and read on the card.

axis Speakers

Who spoke, recruited by dialect region, age and background, so the long tail is represented rather than averaged away.

axis Speech type

The speech a benchmark removes and production does not. The frontier asks for overlap and for what happens between the words.

Where a required range is not in inventory, YPAI says so at qualification and scopes the collection that fills it.

QUALIFICATION

An existing item is reviewed against your use before it is offered.

Six checks, run by a person who has read your brief, on every item YPAI owns or represents. An item that fails one is set aside. An item that passes is offered with its card.

  1. Intended use

    Training, evaluation or a deployed product. The permission differs for each, and the item's licence basis has to carry the one you need.

  2. Speech type

    Read, spontaneous, overlapping, code-switched, expressive. The item's elicitation is read against the speech your system meets in production.

  3. Acoustic conditions

    Setting, device tier and channel, as captured. A studio item is set aside for a cabin deployment and says so.

  4. Format

    Container, channels, transcript schema and metadata levels, checked against what your receiving system reads.

  5. Provenance

    Where the recordings come from, under which agreement, and whether the chain from contributor to licensor is unbroken.

  6. Permitted rights

    What the item may be used for, for how long, where, and whether a derivative model may be deployed. Read on the card, granted by the licence.

a person reviews, then offers

your condition, as described above

select a value on each axis and the request line fills in

ONE DATASET CARD

What one dataset card says. Every field a name, every value a word.

An example card, in the order buyers read it. Under NDA the same card carries the item's own values; chain of title and provenance gaps are shared under NDA only.

A few seconds from your microphone are measured in this tab and discarded. The measurement fills the acoustic fields of the card in words and adds them to the request line.

example card · no item named

Language and dialect
one language, one written norm, the dialect region tagged
Acoustic environment
the setting as captured, noise and reverberation classed
Elicitation
the speech type as prompted or as it happened
Transcript schema
verbatim with disfluencies, speaker turns, word times
Device and channel
the tier it was captured on, one channel per speaker
Age band and gender
recorded as categories per speaker
Native or second language
recorded per speaker
Self-reported or checked
dialect checked by a linguist from the region
Licence basis and permitted uses
commercial training · evaluation only, held out of training · research
Exclusions
voice cloning, resale of raw audio, white-label supply
Consent scope and withdrawal
per speaker, per purpose; withdrawal quarantines the recordings and is reported
Chain of title and provenance
shared under NDA
Derivative-model rights
as the schedule sets them, named before signature
Sublicensing and resale
prohibited unless the schedule says otherwise
Term and territory
as the schedule sets them
Expiry and provenance gaps
shared under NDA

Ships with every item

  • the dataset card, in this order
  • a data card in the order of the EU AI Act's data-governance article
  • a Croissant description carrying provenance and permitted use
  • a consent record per speaker
  • capture metadata per recording
  • the transcript schema and its guidelines
  • the QA trail, per dialect group

THE GATE

Under NDA you read the inventory. The licence grants the use.

The NDA is the first gate and only that. It opens the list. Every use is granted by the licence that follows it.

What the NDA opens

Dataset cards, the inventory, prices, provenance descriptions, technical specifications, anonymised or controlled samples, partner and sourcing information where the agreement allows it.

What only the licence grants

Training rights, evaluation rights on production data, resale, sublicensing, white-label rights, access to personal data, keeping samples permanently.

What follows the NDA

Dataset Licence Agreement, or MSA with a SOW and a Data Licence Schedule.

The licence structures

An evaluation licence, held out of training, a research licence, a commercial training licence, an exclusive structure, with geography, use or term set.

Samples

Samples are for evaluation. YPAI can ask for samples and evaluation material to be deleted.

Who opens it

A named project lead opens the inventory with you.

THE GAP

Where the inventory has the gap, YPAI records it.

An axis with no matching item is collected as new recordings. Recruitment by dialect region, setting and device, on the platform the inventory was built on, so the new recordings arrive with the same card as the licensed ones.

Language and dialect
one language, one written norm, the dialect region tagged
Acoustic environment
the setting as captured, noise and reverberation classed
Elicitation
the speech type as prompted or as it happened
Device and channel
the tier it was captured on, one channel per speaker
open field · collected new

THE BRIEF

Describe the speech you need. Start under NDA.

Two sentences are enough. A named project lead reads it, matches it against own and partner inventory, and comes back with what can be licensed and what has to be recorded.

Speech data Dataset licensing

Speech type (optional)

QUESTIONS

Questions ASR and TTS teams ask

Why is there no list of datasets on this page?

The inventory, the dataset cards, the prices and the provenance descriptions are shared under NDA. The page shows the axes an item is described on and one example card, so you can tell before the NDA whether the shape fits.

What does "YPAI owns or represents" mean?

Some items YPAI recorded and controls with sufficient rights. Others YPAI represents for partner suppliers, after sampling, technical review and a rights check. Both arrive with the same card and the same licence structure.

Can the same item be licensed for training and for evaluation?

Only if the schedule says so. Evaluation, model training and commercial deployment are separate permissions, each with territory, term, withdrawal handling, deletion and retention written in. An evaluation set held out of training stays held out.

What happens when a speaker withdraws consent after delivery?

The recordings are quarantined in YPAI systems, the withdrawal is reported to you with a ledger entry, and erasure follows within the period the schedule sets.

What if no item covers our condition?

YPAI says so at qualification and scopes the collection that fills the gap: recruitment by dialect region, setting and device, on the same platform, delivered with the same card.

Where is the audio processed and stored?

Primary operations are in Norway on EEA infrastructure. Items can be structured for EEA-only sourcing and EEA-resident processing; other delivery is arranged where a deployment requires it.

Described by the conditions it was recorded in.

Describe the condition. The card that matches comes back under NDA; the field that does not is collected as new recordings.

Describe the speech you need