Speech data · Technical specifications

Formats, rates, manifests and QA outputs. The spec sheet your pipeline reads.

WAV or FLAC, 16 to 48 kHz, 16 or 24-bit, structured JSON manifests, per-recording QA status. The project spec sheet is finalized during scoping.

Speech dataset · specification

Format
WAV · FLAC
Sample rate
16 · 44.1 · 48 kHz
Bit depth
16 · 24-bit
Channels
Mono · stereo on request
Capture
48 kHz / 24-bit
Signal-to-noise
40 dB or better
Segmentation
Utterance
Manifest
JSON, per recording
Integrity
SHA-256 per file
Description
Croissant 1.1 + data card
Acceptance
98% first-delivery floor
Pilot
within 10 business days

Project sheet finalized during scoping · thresholds per project

The numbers capture · acceptance

Four numbers the specification carries, measured per recording and reported per delivery.

Detailed specs vary by project. A finalized spec sheet is provided during scoping.

Capture floor · acceptance floor · quota tolerance

  • 48 kHz Capture sample rate 24-bit, studio and field
  • 40 dB Signal-to-noise or better, measured per recording
  • 5 pts Quota tolerance per cell, reported each delivery
  • 98% First-delivery acceptance floor, re-collected below it

Audio standards capture · validation

Constraints are enforced at capture. Invalid audio never enters the pipeline.

Specific thresholds (SNR, noise floor dB, environment requirements) are defined per project during technical scoping.

Gates, in capture order

  1. Device Device class enforced at capture time
  2. Sample rate Matches the project specification before submission
  3. Noise floor Within the agreed threshold, checked automatically
  4. Clipping No clipping or distortion detected
  5. Duration Within the expected range for the prompt
  6. Rejection Non-conforming audio rejected before it enters the pipeline

Recording constraints: Device constraints enforced at capture time, sample rate validation before submission, noise floor checks applied automatically, invalid audio rejected before entering pipeline.

Validation checks: Sample rate matches project specification, environment noise floor within acceptable threshold, no clipping or distortion detected, audio duration within expected range.

Annotation and segmentation utterance · transcript · manifest

Segmented at the utterance, transcribed to your convention, described in the manifest.

Complete schema documentation is provided during scoping. Additional fields available on request.

Segmentation approach: Recordings are segmented at the utterance level by default. File-level segmentation or alternative approaches are available on request and specified during project scoping.

Transcript and label formats: Transcripts are delivered as verbatim or normalized text depending on project requirements. Label formats are aligned with common training pipeline conventions.

Manifest schema: Datasets include structured JSON manifests with fields such as recording_id, speaker_id, transcript, duration_ms, sample_rate, format, language, region, consent_reference, and qa_status.

manifest.json · one record, example values
{"recording_id": "rec_004217","speaker_id": "spk_0142","transcript": "...","duration_ms": 4210,"sample_rate": 48000,"format": "wav","language": "nb-NO","region": "Vestlandsk","consent_reference": "cns_8f31","qa_status": "accepted"}

Quality assurance automated · human · dataset-level

Three QA stages, each one leaving an output in the delivery.

  1. Automated validation

    SNR, clipping and silence thresholds, synthetic-artifact detection, specification compliance, automatic rejection of non-conforming recordings.

    Output Per-recording validation result in the manifest

  2. Human QA

    Review coverage follows the agreed acceptance and sampling plan, with linguistic correctness, naturalness and script adherence where applicable.

    Output Per-recording acceptance decision

  3. Dataset-level QA

    Coverage balance, speaker distribution, label integrity, final acceptance review before delivery.

    Output QA report with the delivery

Delivery and handoff transfer · versioning · support

File-based delivery, versioned, with the handoff procedure in the agreement.

Handoff procedures, acceptance criteria, and post-delivery support are documented in the project agreement.

Delivery method: Delivery methods are defined during scoping and may include secure transfer, cloud storage handoff, or other enterprise-compatible mechanisms.

Versioning and iterations: Iteration cycles, revision policies, and version control are defined during project scoping and are contract-bound.

Integration wav · flac · json

Delivered for offline training. Folder structure and naming documented per project.

YPAI datasets are delivered in formats compatible with standard ML training pipelines. Delivery is file-based and designed for offline training workflows.

Delivery package, structure documented per project
  • dataset_v1/
  • manifest.json
  • croissant.json
  • data_card.md
  • qa_report.json
  • audio/
  • spk_0142/
  • rec_004217.wav
  • transcripts/
  • rec_004217.txt
  • checksums.sha256

Integration overview: Datasets delivered in WAV/FLAC with structured JSON manifests, folder structure and naming conventions documented per project, compatible with common ASR/TTS training frameworks. Integration spec available on request during scoping.

What this page covers conventions · qa · delivery

This page describes the conventions. The project specification is written in scoping.

This page describes

  • Typical formats and conventions
  • Standard QA process outputs
  • Delivery and integration overview

Written elsewhere

  • Project-specific specifications, finalized during scoping
  • Pricing and commercial terms
  • Procurement appendices, legal terms and DPA documentation, linked from the main Speech Data page or provided during enterprise consultation

Questions

Formats, metadata and delivery

What audio formats does YPAI support for speech datasets?

YPAI delivers speech datasets in WAV and FLAC formats. The specific format is defined during project scoping based on your pipeline requirements.

What sample rates are available?

Standard sample rates include 16 kHz, 44.1 kHz, and 48 kHz. The appropriate sample rate for your project is determined during technical scoping based on your use case and training requirements.

How is audio quality validated?

YPAI combines automated validation (SNR checks, clipping detection, and technical compliance) with Human QA applied according to the agreed acceptance and sampling plan. The plan defines which recordings receive review before acceptance.

What metadata is included with delivered datasets?

Datasets include structured JSON manifests containing fields such as recording_id, speaker_id, transcript, duration_ms, sample_rate, format, language, region, consent_reference, and qa_status. Additional fields are available on request.

Can I request stereo recordings instead of mono?

Yes. The default configuration is mono, but stereo recordings are available on request and can be specified during project scoping.

How are transcripts formatted?

Transcripts are delivered as verbatim or normalized text depending on project requirements. The specific format and conventions are documented during scoping.

What segmentation approach is used?

Recordings are segmented at the utterance level by default. File-level segmentation or alternative approaches are available on request and specified during project scoping.

How is data delivered?

Delivery methods are defined during scoping and may include secure transfer, cloud storage handoff, or other enterprise-compatible mechanisms. Specific options are documented in the project agreement.

Are YPAI datasets compatible with standard ML frameworks?

Yes. Datasets are delivered in formats compatible with standard ML training pipelines, including common ASR and TTS training frameworks. Integration specifications are available on request during scoping.

What QA documentation is provided with delivered datasets?

Delivered datasets include QA reports documenting validation results, coverage balance verification, speaker distribution analysis, and label integrity checks. Specific documentation scope is defined during project scoping.

Can specifications be customized for my project?

Yes. Technical specifications are defined per engagement and finalized during scoping. YPAI works with your team to define project-specific requirements, thresholds, and deliverables.

Is there an API for accessing datasets?

YPAI datasets are delivered as file-based packages designed for offline training workflows. Delivery is by secure transfer or cloud storage handoff, agreed in scoping.

Bring the pipeline requirement. The spec sheet comes back from scoping.

A scoped, confidential discussion with the data team defines the project-specific thresholds, formats and deliverables.