Voice recording for TTS and voice AI. One voice, cast, directed and licensed in writing.
YPAI casts the voice, directs the sessions and writes the rights down before the first take. The model learns from recordings that carry the performance you specified and the permissions you need.
Anyone can record a voice. The rights decide what you can build.
A TTS, cloning or assistant voice needs more than a clean file. It needs to carry:
A voice file
- a recording in a file format
A contracted voice
- who the voice is, chosen against the brief
- the pronunciation reference it was checked against
- the register, pace and emotion that were directed
- the room, microphone and format held constant
- every take, logged and reviewed
- the script it was written for
- the talent's documented consent to the use
- the training and evaluation rights
- the cloning, exclusivity and resale grants
- the term and territory the grants run for
YPAI delivers these as one record around the voice, with the licence written before the session.
The same holds when the voice already exists: a recording becomes usable for a model when its casting, session and rights can be shown in writing. If the gap is coverage across speakers and conditions, that is a managed collection, on its own route.
The project is built around the use the voice must serve.
That use may be:
- a product voice for a TTS system
- a cloned voice, with the talent's documented consent
- an assistant persona across languages
- narration for long-form content
- a brand voice under exclusivity
The casting, session plan, take log and licence schedule follow that use.
Start with the voice you need.
YPAI can run one defined stage or the whole recording programme.
-
Entry point: Cast and record a new voice
Start here when: You need one voice that your data does not contain.
Typical output Recorded voice files, directed take by take, with the take log and the licence schedule.
-
Entry point: Extend a voice you already use
Start here when: A model needs more of an existing voice: pickups, new registers, new scripts.
Typical output Additional sessions in the same chain, logged against the original record.
-
Entry point: Collect a population instead
Start here when: The gap is coverage across speakers and conditions.
Typical output A managed collection against agreed speakers, conditions and checks, on its own route.
CASTING
The voice is chosen against the brief, not from a roster.
The casting brief fixes language, accent, register, age range and the pronunciation reference. Candidates record the same passage under the same conditions, and the choice is made on those takes.
Brief → same passage, same conditions → one chosen
The chosen voice's consent to the intended use is documented before production.
The rights are written down before the first take.
The exact schedule is defined for the customer's use, model and territory.
A voice licence schedule can include the following grants.
Example licence schedule
- Permitted use
-
- model training
- fine-tuning
- evaluation
- regression sets
- inference
- internal testing
- derivative models
- permitted products
- Voice identity
-
- voice cloning
- documented consent to cloning
- likeness
- name and attribution
- synthetic re-use
- style transfer
- Commercial scope
-
- buyout
- exclusivity
- resale
- sublicensing
- distribution
- white-label
- end-client disclosure
- reseller terms
- Term and territory
-
- term
- territory
- renewal
- termination
- post-term use
- archive
- Session record
-
- casting brief
- pronunciation reference
- script version
- take log
- reviewer decision
- delivered files
- format
- Provenance and indemnity
-
- consent record
- rights status
- indemnity
- deletion right
- audit trail
- signed schedule
- schedule version
The licence schedule is signed before the first take.
Voice-cloning, buyout, exclusivity, resale, sublicensing, distribution and white-label rights need an explicit written grant.
Every take is
directed, checked
and named.
The capture chain is fixed before the first take: room, microphone, format and the script version.
Register, pace and emotion,
directed take by take.
A director holds register, pace and emotion constant across takes, and pickups are recorded in the same chain.
- Brief
- Casting
- Direction
- Takes
- Review
The session plan is built against the customer's script, use and pronunciation reference.
Studio-clean is one condition.
The brief sets the others.
Train a voice for a car, a call centre or a kitchen and the booth is the wrong room. Conditions are set per brief; microphone and format are held constant inside each.
The room The chain - Booth
- Office
- Vehicle
- Home
- Street
- Café
The room is part of the brief; the microphone and the format stay fixed inside it.
Delivery is defined before the session, not after.
Every take is accepted or sent back before the files are named and delivered.
Files are delivered as WAV or FLAC at 48 kHz with the take log, the script alignment and the licence schedule. Noise floor, naming and transfer follow the technical specification agreed for the project.
Ships with the accepted takes / take log · script alignment · pronunciation reference · session record · consent record · licence schedule · checksums
Delivered as WAV or FLAC, or mapped to the customer's TTS training layout. Format support is checked against the actual toolchain and delivery requirement.
Different voices need different sessions. TTS product voice · cloned voice with documented consent · assistant persona · narration · brand voice under exclusivity. Cast from contributors in 150+ languages.
QUESTIONS
Questions TTS and voice AI teams ask
Which rights go into the licence schedule?
The schedule is defined for your use, model and territory. It can grant model training, fine-tuning and evaluation, voice cloning with the talent's documented consent, buyout, exclusivity, resale and sublicensing, and it sets the term and territory. It is signed before the first take.
How is the voice chosen?
The casting brief fixes language, accent, register, age range and the pronunciation reference. Candidates record the same passage under the same conditions, and the voice is chosen on those takes. The chosen voice's consent to the intended use is documented before production.
Can sessions be recorded outside a studio?
Yes. The brief sets the room: booth, office, vehicle, home, street or café. The microphone and the format are held constant inside each condition.
What is delivered with the recordings?
WAV or FLAC at 48 kHz, or files mapped to your TTS training layout, with the take log, script alignment, pronunciation reference, session record, consent record, licence schedule and checksums. Noise floor, naming and transfer follow the technical specification agreed for the project.
Can you extend a voice we already use?
Yes. Pickups, new registers and new scripts are recorded in the same chain and logged against the original record. When the gap is coverage across many speakers and conditions, the project is a managed collection instead.
Which languages can you cast in?
Voices are cast from contributors in 150+ languages.
Send the script, the voice brief and the rights you need.
The use, the language and register, the script, the conditions, the rights and the delivery format. YPAI returns the casting plan, the session plan and the licence schedule to sign off before the first take.
Related
- Speech data The hub: collection, voice recording, annotation, evaluation and licensing.
- Transcription and labels Transcription, diarisation, alignment and linguistic annotation.
- Dataset review and licensing Existing or partner-sourced datasets against intended use and rights.
- Coverage and recruitment Feasibility by language, country, accent and cohort.
- Technical specifications Formats, sample rates, noise floor and delivery.
- Evaluation sets Rubrics, error taxonomy and regression sets for a voice model.
- Pilots Validate casting, session and acceptance before scale.