A population teaches a model to listen. One voice teaches it to speak.
YPAI casts the voice, directs the sessions and writes the rights down before the first take. The model learns from recordings that carry the performance you specified and the permissions you need.
Anyone can record a voice. The rights decide what you can build.
A TTS, cloning or assistant voice needs more than a clean file. It needs to carry:
A voice file
- a recording
- a file format
A contracted voice
- who the voice is, chosen against the brief
- the pronunciation reference it was checked against
- the register, pace and emotion that were directed
- the room, microphone and format held constant
- every take, logged and reviewed
- the script it was written for
- the talent's documented consent to the use
- the training and evaluation rights
- the cloning, exclusivity and resale grants
- the term and territory the grants run for
YPAI delivers these as one record around the voice, with the licence written before the session rather than attached to the files afterwards.
The same holds when the voice already exists: a recording becomes usable for a model when its casting, session and rights can be shown in writing. If the gap is coverage across speakers and conditions, that is a managed collection, on its own route.
The project is built around the use the voice must serve.
That use may be:
- a product voice for a TTS system
- a cloned voice, with the talent's documented consent
- an assistant persona across languages
- narration for long-form content
- a brand voice under exclusivity
Thecasting,session plan,take logandlicence schedulefollow that use.
Start with the voice you need.
YPAI can run one defined stage or the whole recording programme.
-
Entry point: Cast and record a new voice
Start here when: You need one voice that your data does not contain.
Typical output Recorded voice files, directed take by take, with the take log and the licence schedule.
-
Entry point: Extend a voice you already use
Start here when: A model needs more of an existing voice: pickups, new registers, new scripts.
Typical output Additional sessions in the same chain, logged against the original record.
-
Entry point: Collect a population instead
Start here when: The gap is coverage across speakers and conditions, not one voice.
Typical output A managed collection against agreed speakers, conditions and checks, on its own route.
CASTING
The voice is chosen against the brief, not from a roster.
The casting brief fixes language, accent, register, age range and the pronunciation reference. Candidates record the same passage under the same conditions, and the choice is made on those takes.
Brief → same passage, same conditions → one chosen
The chosen voice's consent to the intended use is documented before production.
The rights are written down before the first take.
The exact schedule is defined for the customer's use, model and territory.
A voice licence schedule can include the following grants.
Representative licence schedule
- Permitted use
-
- model training
- fine-tuning
- evaluation
- regression sets
- inference
- internal testing
- derivative models
- permitted products
- Voice identity
-
- voice cloning
- documented consent to cloning
- likeness
- name and attribution
- synthetic re-use
- style transfer
- Commercial scope
-
- buyout
- exclusivity
- resale
- sublicensing
- distribution
- white-label
- end-client disclosure
- reseller terms
- Term and territory
-
- term
- territory
- renewal
- termination
- post-term use
- archive
- Session record
-
- casting brief
- pronunciation reference
- script version
- take log
- reviewer decision
- delivered files
- format
- Provenance and indemnity
-
- consent record
- rights status
- indemnity
- deletion right
- audit trail
- signed schedule
- schedule version
The licence schedule is signed before the first take.
Voice-cloning, buyout, exclusivity, resale, sublicensing, distribution and white-label rights need an explicit written grant.
Every take is
directed,
checked, named.
The capture chain is fixed before the first take: room, microphone, format and the script version.
Direct the take,
not only record it.
A director holds register, pace and emotion constant across takes, and pickups are recorded in the same chain.
- Brief
- Casting
- Direction
- Takes
- Review
The session plan is built against the customer's script, use and pronunciation reference.
Studio-clean is one condition,
not the default.
Train a voice for a car, a call centre or a kitchen and the booth is the wrong room. Conditions are set per brief; microphone and format are held constant inside each.
The room The chain - Booth
- Office
- Vehicle
- Home
- Street
- Café
The room is part of the brief; the microphone and the format stay fixed inside it.
Delivery is defined before the session, not after.
Every take is accepted or sent back before the files are named and delivered.
Files are delivered as WAV or FLAC at 48 kHz with the take log, the script alignment and the licence schedule. Noise floor, naming and transfer follow the technical specification agreed for the project.
Ships with the accepted takes / take log · script alignment · pronunciation reference · session record · consent record · licence schedule · checksums
Delivered as WAV or FLAC, or mapped to the customer's TTS training layout. Format support is confirmed against the actual toolchain and delivery requirement.
Different voices need different sessions. TTS product voice · cloned voice with documented consent · assistant persona · narration · brand voice under exclusivity. Cast from contributors in 150+ languages.
Send the script, the voice brief and the rights you need.
The use, the language and register, the script, the conditions, the rights and the delivery format. YPAI returns the casting plan, the session plan and the licence schedule to sign off before the first take.
Related
- Speech data The hub: collection, voice recording, annotation, evaluation and licensing.
- Managed collection Population coverage across speakers, conditions and scenarios.
- Transcription and labels Transcription, diarisation, alignment and linguistic annotation.
- Dataset review and licensing Existing or partner-sourced datasets against intended use and rights.
- Coverage and recruitment Feasibility by language, country, accent and cohort.
- Technical specifications Formats, sample rates, noise floor and delivery.
- Evaluation sets Rubrics, error taxonomy and regression sets for a voice model.
- Pilots Validate casting, session and acceptance before scale.