A voice talent in profile at a condenser microphone in a vocal booth, headphones on, script in hand, seen through the control-room glass

A population teaches a model to listen. One voice teaches it to speak.

YPAI casts the voice, directs the sessions and writes the rights down before the first take. The model learns from recordings that carry the performance you specified and the permissions you need.

Representative session
First take Pickup Reviewed Accepted take

Anyone can record a voice. The rights decide what you can build.

A TTS, cloning or assistant voice needs more than a clean file. It needs to carry:

A voice file

  • a recording
  • a file format

A contracted voice

  • who the voice is, chosen against the brief
  • the pronunciation reference it was checked against
  • the register, pace and emotion that were directed
  • the room, microphone and format held constant
  • every take, logged and reviewed
  • the script it was written for
  • the talent's documented consent to the use
  • the training and evaluation rights
  • the cloning, exclusivity and resale grants
  • the term and territory the grants run for

YPAI delivers these as one record around the voice, with the licence written before the session rather than attached to the files afterwards.

The same holds when the voice already exists: a recording becomes usable for a model when its casting, session and rights can be shown in writing. If the gap is coverage across speakers and conditions, that is a managed collection, on its own route.

The project is built around the use the voice must serve.

That use may be:

  • a product voice for a TTS system
  • a cloned voice, with the talent's documented consent
  • an assistant persona across languages
  • narration for long-form content
  • a brand voice under exclusivity

Thecasting,session plan,take logandlicence schedulefollow that use.

Start with the voice you need.

YPAI can run one defined stage or the whole recording programme.

  1. Entry point: Cast and record a new voice

    Start here when: You need one voice that your data does not contain.

    Typical output Recorded voice files, directed take by take, with the take log and the licence schedule.

  2. Entry point: Extend a voice you already use

    Start here when: A model needs more of an existing voice: pickups, new registers, new scripts.

    Typical output Additional sessions in the same chain, logged against the original record.

  3. Entry point: Collect a population instead

    Start here when: The gap is coverage across speakers and conditions, not one voice.

    Typical output A managed collection against agreed speakers, conditions and checks, on its own route.

A director from behind wearing headphones, reading a fanned stack of candidate sheets beside a laptop with a soft waveform glow

CASTING

The voice is chosen against the brief, not from a roster.

The casting brief fixes language, accent, register, age range and the pronunciation reference. Candidates record the same passage under the same conditions, and the choice is made on those takes.

Brief  →  same passage, same conditions  →  one chosen

The chosen voice's consent to the intended use is documented before production.

The rights are written down before the first take.

The exact schedule is defined for the customer's use, model and territory.

A voice licence schedule can include the following grants.

Representative licence schedule

Permitted use
  • model training
  • fine-tuning
  • evaluation
  • regression sets
  • inference
  • internal testing
  • derivative models
  • permitted products
Voice identity
  • voice cloning
  • documented consent to cloning
  • likeness
  • name and attribution
  • synthetic re-use
  • style transfer
Commercial scope
  • buyout
  • exclusivity
  • resale
  • sublicensing
  • distribution
  • white-label
  • end-client disclosure
  • reseller terms
Term and territory
  • term
  • territory
  • renewal
  • termination
  • post-term use
  • archive
Session record
  • casting brief
  • pronunciation reference
  • script version
  • take log
  • reviewer decision
  • delivered files
  • format
Provenance and indemnity
  • consent record
  • rights status
  • indemnity
  • deletion right
  • audit trail
  • signed schedule
  • schedule version

The licence schedule is signed before the first take.

Voice-cloning, buyout, exclusivity, resale, sublicensing, distribution and white-label rights need an explicit written grant.

A sound engineer from behind at the console, a condenser microphone and pop filter in front, the voice talent visible through the glass.
Session

Every take is
directed,
checked, named.

The capture chain is fixed before the first take: room, microphone, format and the script version.

A vocal booth with a condenser microphone on a boom arm, a pop filter, and a script on a music stand.
Script and pronunciation reference
A condenser microphone with a reflection filter, headphones and a lamp on a desk.
One chain, held constant

Direct the take,
not only record it.

A director holds register, pace and emotion constant across takes, and pickups are recorded in the same chain.

A vocal booth, condenser microphone and script stand under one lamp.
  1. Brief
  2. Casting
  3. Direction
  4. Takes
  5. Review
Take / logged with script line and numberReview / reviewer decision recordedPickup / same chain, same day where the schedule allowsAccepted / named and delivered

The session plan is built against the customer's script, use and pronunciation reference.

Studio-clean is one condition,
not the default.

Train a voice for a car, a call centre or a kitchen and the booth is the wrong room. Conditions are set per brief; microphone and format are held constant inside each.

A condenser microphone with a reflection filter on an office desk at night, headphones, a laptop and blinds behind. The room The chain
  1. Booth
  2. Office
  3. Vehicle
  4. Home
  5. Street
  6. Café

The room is part of the brief; the microphone and the format stay fixed inside it.

A fountain pen resting on the signature line of a signed voice rights agreement.
Ships with the signed schedule

Delivery is defined before the session, not after.

Every take is accepted or sent back before the files are named and delivered.

AcceptedPickupRe-recordedReplacedRejected

Files are delivered as WAV or FLAC at 48 kHz with the take log, the script alignment and the licence schedule. Noise floor, naming and transfer follow the technical specification agreed for the project.

Ships with the accepted takes  /  take log · script alignment · pronunciation reference · session record · consent record · licence schedule · checksums

Delivered as WAV or FLAC, or mapped to the customer's TTS training layout. Format support is confirmed against the actual toolchain and delivery requirement.

Different voices need different sessions.  TTS product voice · cloned voice with documented consent · assistant persona · narration · brand voice under exclusivity. Cast from contributors in 150+ languages.

Send the script, the voice brief and the rights you need.

The use, the language and register, the script, the conditions, the rights and the delivery format. YPAI returns the casting plan, the session plan and the licence schedule to sign off before the first take.

Related

Scope a voice recording

The form adapts to the work, asks only for relevant details and sends your brief to the person who can act on it.

Service required

Your selection routes the brief to the right person.

A named project lead reviews every enquiry and replies within one business day