---
title: "Voice Recording for TTS and Voice AI | YPAI"
url: https://ypai.ai/audio/voice-recording/
description: "Casting, directed recording sessions and model-use rights in writing for TTS, voice-cloning and assistant voices. One voice, under contract."
source: "src/copy/routes (route /audio/voice-recording/)"
---

# Voice recording for TTS and voice AI. One voice, cast, directed and licensed in writing.

> Casting, directed recording sessions and model-use rights in writing for TTS, voice-cloning and assistant voices. One voice, under contract.

YPAI casts the voice, directs the sessions and writes the rights down before the first take. The model learns from recordings that carry the performance you specified and the permissions you need.

## Anyone can record a voice. The rights decide what you can build.

A TTS, cloning or assistant voice needs more than a clean file. It needs to carry:

- a recording in a file format
- who the voice is, chosen against the brief
- the pronunciation reference it was checked against
- the register, pace and emotion that were directed
- the room, microphone and format held constant
- the script it was written for
- the talent's documented consent to the use
- the cloning, exclusivity and resale grants
- the term and territory the grants run for

YPAI delivers these as one record around the voice, with the licence written before the session.

The same holds when the voice already exists: a recording becomes usable for a model when its casting, session and rights can be shown in writing. If the gap is coverage across speakers and conditions, that is a {link}, on its own route.

The project is built around the use the voice must serve.

- a product voice for a TTS system
- a cloned voice, with the talent's documented consent

### Start with the voice you need.

YPAI can run one defined stage or the whole recording programme.

#### Cast and record a new voice

- You need one voice that your data does not contain.
- Recorded voice files, directed take by take, with the take log and the licence schedule.

#### Extend a voice you already use

- A model needs more of an existing voice: pickups, new registers, new scripts.
- Additional sessions in the same chain, logged against the original record.

#### Collect a population instead

- The gap is coverage across speakers and conditions.
- A managed collection against agreed speakers, conditions and checks, on its own route.

## The voice is chosen against the brief, not from a roster.

A director from behind, headphones on, reading a fanned stack of candidate sheets beside a laptop with a soft waveform glow

The casting brief fixes language, accent, register, age range and the pronunciation reference. Candidates record the same passage under the same conditions, and the choice is made on those takes.

Brief, same passage under the same conditions, one chosen

Brief&nbsp; &#8594; &nbsp;same passage, same conditions&nbsp; &#8594; &nbsp;one chosen

The chosen voice's consent to the intended use is documented before production.

## The rights are written down before the first take.

The exact schedule is defined for the customer's use, model and territory.

A voice licence schedule can include the following grants.

### Permitted use

### Voice identity

### Commercial scope

### Term and territory

### Session record

### Provenance and indemnity

The licence schedule is signed before the first take.

Voice-cloning, buyout, exclusivity, resale, sublicensing, distribution and white-label rights need an explicit

The capture chain is fixed before the first take: room, microphone, format and the script version.

A director holds register, pace and emotion constant across takes, and pickups are recorded in the same chain.

- Take / logged with script line and number
- Pickup / same chain, same day where the schedule allows

The session plan is built against the customer's script, use and pronunciation reference.

Train a voice for a car, a call centre or a kitchen and the booth is the wrong room. Conditions are set per brief; microphone and format are held constant inside each.

The room is part of the brief; the microphone and the format stay fixed inside it.

## Delivery is defined before the session, not after.

Every take is accepted or sent back before the files are named and delivered.

Files are delivered as WAV or FLAC at 48 kHz with the take log, the script alignment and the licence schedule. Noise floor, naming and transfer follow the technical specification agreed for the project.

- Ships with the accepted takes &nbsp;<em>/</em>&nbsp; take log &middot; script alignment &middot; pronunciation reference &middot; session record &middot; consent record &middot; licence schedule &middot; checksums
- Delivered as WAV or FLAC, or mapped to the customer's TTS training layout. Format support is checked against the actual toolchain and delivery requirement.

<b>Different voices need different sessions.</b> &nbsp;TTS product voice &middot; cloned voice with documented consent &middot; assistant persona &middot; narration &middot; brand voice under exclusivity. Cast from contributors in 150+ languages.

## Related

### [Speech data](https://ypai.ai/speech-data/)

- The hub: collection, voice recording, annotation, evaluation and licensing.

### [Transcription and labels](https://ypai.ai/annotation/audio-speech-annotation-services/)

- Transcription, diarisation, alignment and linguistic annotation.

### [Dataset review and licensing](https://ypai.ai/audio/datasets/)

- Existing or partner-sourced datasets against intended use and rights.

### [Coverage and recruitment](https://ypai.ai/speech-data/language-coverage/)

- Feasibility by language, country, accent and cohort.

### [Technical specifications](https://ypai.ai/speech-data/technical-specifications/)

- Formats, sample rates, noise floor and delivery.

### [Evaluation sets](https://ypai.ai/speech-data/evaluation-program/)

- Rubrics, error taxonomy and regression sets for a voice model.

### [Pilots](https://ypai.ai/pilots/)

- Validate casting, session and acceptance before scale.

## Questions TTS and voice AI teams ask

### Which rights go into the licence schedule?

The schedule is defined for your use, model and territory. It can grant model training, fine-tuning and evaluation, voice cloning with the talent's documented consent, buyout, exclusivity, resale and sublicensing, and it sets the term and territory. It is signed before the first take.

### How is the voice chosen?

The casting brief fixes language, accent, register, age range and the pronunciation reference. Candidates record the same passage under the same conditions, and the voice is chosen on those takes. The chosen voice's consent to the intended use is documented before production.

### Can sessions be recorded outside a studio?

Yes. The brief sets the room: booth, office, vehicle, home, street or café. The microphone and the format are held constant inside each condition.

### What is delivered with the recordings?

WAV or FLAC at 48 kHz, or files mapped to your TTS training layout, with the take log, script alignment, pronunciation reference, session record, consent record, licence schedule and checksums. Noise floor, naming and transfer follow the technical specification agreed for the project.

### Can you extend a voice we already use?

Yes. Pickups, new registers and new scripts are recorded in the same chain and logged against the original record. When the gap is coverage across many speakers and conditions, the project is a managed collection instead.

### Which languages can you cast in?

Voices are cast from contributors in 150+ languages.
