Speech data · GDPR-compliant speech data

Voice is personal data, and sometimes biometric. Every second of delivered speech carries its lawful basis.

Speech datasets collected under GDPR Article 6 consent, with Article 7 conditions met, explicit consent where Article 9 applies, and the consent id, timestamp and origin on every file.

EU AI Act obligations for high-risk systems are covered on the AI Act risk classification page.

The article register, and what each one means for the corpus
  1. Art. 6(1)(a) Lawful basis Consent, obtained as its own affirmative action before recording. The basis is defined per engagement and documented.
  2. Art. 7 Conditions for consent Freely given, specific, informed, unambiguous. Recorded with timestamp and version. Withdrawable at any time.
  3. Art. 9 Special categories Where voice is processed for identification or verification it is biometric data, and consent is explicit.
  4. Arts. 12 to 23 Data subject rights Access, rectification, erasure, restriction and objection, handled operationally against the recordings a request touches.
  5. Art. 28 Processor The DPA, executed before production collection. Processor for custom collection, independent controller for licensing.

Compliance overview basis · residency · sourcing

Three facts hold across every dataset. Consent, EU sourcing, a closed collection model.

Every speaker is a contracted, identity-verified contributor recording inside YPAI's own platform. The dataset arrives with the consent chain that makes it usable in production.

GDPR Art. 6 basis
Consent, Art. 6(1)(a)
Biometric derogation
Explicit consent where Art. 9 applies
Data residency
EEA by default for European engagements
Sourcing
Contracted, identity-verified contributors on the YPAI platform
Rights
Access, rectification, erasure, restriction and objection, handled operationally
Lineage
Consent id, collection timestamp and origin log per file

Regulatory exposure art. 6 · art. 9 · lifecycle

Provenance is a go or no-go criterion. Three places speech data fails review.

For enterprise organizations, data provenance is now a go or no-go criterion for model deployment.

GDPR Art. 6
Voice is a biological identifier. Even without metadata, speech content and acoustic markers can re-identify individuals. Treating speech as anonymous by default is legally indefensible; it requires pseudonymisation plus a lawful basis.
GDPR Art. 9
Voice data can constitute biometric data when processed for identification or verification. That triggers special-category status, and the legal bar rises from legitimate interest to explicit consent. Scraped data fails here completely.
Lifecycle
Liability persists across model versions. If the original dataset provenance cannot be verified years later, the model itself is compromised. Snapshot compliance is not enough; audit traceability has to hold for the full retention period.

European regulators are actively scrutinizing the lawful basis of acquisition for training datasets under GDPR Article 6. The era of indiscriminate data scraping is ending.

Legal risk doctrine. Using non-compliant speech data creates a toxic asset. Under the "fruit of the poisonous tree" doctrine, a model trained on illicit data may face mandatory deletion orders.

What YPAI provides datasets · collection · artefacts

Production-grade speech datasets for enterprises that need regulatory certainty.

Off-the-shelf corpora and custom scoped collection, each delivered with the compliance artefacts your legal and procurement teams will ask for.

Core deliverables

Off-the-shelf speech datasets.
Ready-to-deploy libraries of EU-sourced speech for training, evaluation and fine-tuning, delivered as structured, annotated audio corpora with full demographic metadata.
Custom scoped collection.
Rapid execution of specific demographic, acoustic or linguistic requirements. Data is collected within YPAI's controlled platform by contracted contributors.

Included compliance artefacts

Audit documentation package.
Every dataset delivery includes lineage records: consent ids, verified collection timestamps and geographic origin logs for legal defence.
Enterprise engagement models.
Commercial frameworks designed for procurement: Master Services Agreements, Data Processing Agreements and explicit indemnification clauses.

These controls apply to all speech datasets delivered by YPAI.

Lawful basis and consent art. 6 · art. 7 · recital 32

Consent, Article 6(1)(a), backed by an affirmative action from every speaker.

YPAI structures all data collection under defined GDPR Article 6 lawful bases, primarily consent. Controlled collection means every second of speech in a delivered dataset carries an affirmative action from the data subject.

Consent must be freely given, specific, informed, and unambiguous indication of the data subject's wishes.

GDPR Recital 32

Detailed consent framework documentation, including exact copies of the participant agreements used for your dataset, is available for legal review during evaluation.

Data subject rights articles 12 to 23

Rights are handled in operational terms. A request maps to the recordings it touches.

Right of access.
Indexed metadata is maintained for all contributors. On an authenticated subject access request, the repository is queried to locate the specific recordings associated with a user id within any delivered dataset.
Right to erasure.
When a valid deletion request is processed, data is purged from active storage. "Do not use" flags propagate to client deliverables where contractually enforceable. Where deletion is not technically reversible in trained models, YPAI documents the withdrawal and enforces non-use in future training and deliveries, consistent with prevailing regulatory guidance. Deletion logs are kept to prove compliance during audits.

This framework governs how YPAI delivers speech data into production AI environments, so the defensibility holds for the model's lifetime.

Provenance and audit lineage · per file

Having data and defending data differ by one thing, provenance.

YPAI speech datasets are constructed assets with complete lineage. Every audio file delivered links to a specific collection event, a verified contributor profile and a timestamped consent record.

Long-term auditability. The metadata structure lets a client answer audit questions years after deployment: where did this specific training vector come from, and did we have the right to use it?

Metadata structure, JSON-LD compatible
{
  "file_id": "ypai_v4_29841",
  "origin": "EU_FR_PARIS",
  "consent_id": "c_9928_v2_signed",
  "lawful_basis": "GDPR_ART_6_1_A",
  "demographics": {
    "yob": 1992,
    "gender": "female"
  },
  "collection_date": "2024-02-14T10:00:00Z"
}

Residency, roles and security eea · msa · toms

Where the data lives, who controls it, and how it is protected.

Cloud regions
Frankfurt, Dublin (AWS and GCP)
Transfers
Governed by SCCs where applicable
Custom collection
YPAI as data processor
Licensing
YPAI as independent controller
Encryption
AES-256 at rest, strict RBAC
Retention
Automated deletion at end of defined purpose
Data residency and sovereignty.
By default, YPAI processes and stores data within the EEA for European engagements. Cloud regions are Frankfurt and Dublin (AWS and GCP); transfers are governed by SCCs where applicable.
Controller and processor.
Roles are explicitly defined in the Master Services Agreement and follow the engagement structure: data processor for custom collection, independent controller for licensing.
Security and retention.
Technical and organisational measures include AES-256 encryption at rest and strict role-based access control, with automated retention and deletion policies so data is held only for its defined purpose.

Closed-loop collection contracted · verified · direct

Contracted contributors under a direct legal relationship. The provenance that open marketplaces cannot reconstruct.

YPAI runs a closed-loop collection service. Contributors are contracted and identity-verified, and the legal relationship with each data subject is direct.

Why open crowdsourcing fails review

  • True speaker identity cannot be verified, which opens the door to Sybil attacks
  • Farmed or synthetic data is easily injected
  • Consent is weakly enforceable across jurisdictions

What the closed loop gives you

  • Contracted contributors with verified identities
  • Device fingerprinting and environment checks on every session
  • A direct legal relationship with every data subject

Request one business day

Bring the requirement. The compliance documentation for your jurisdiction follows.

Tell us the use case, the timeline and the compliance requirement. Our team replies within one business day with the documentation relevant to your case.

Legal and procurement

Basis, biometrics and rights

Which lawful basis applies to a delivered dataset?

Consent under GDPR Article 6(1)(a) is the primary basis. Where voice is processed for identification or verification and Article 9 applies, consent is explicit. The applicable basis is defined per engagement and documented in the participant materials and the DPA.

How is consent demonstrated during an audit?

Every delivered file carries a consent id, a verified collection timestamp and an origin log. The exact participant agreements used for your dataset are available for legal review during evaluation.

What happens when a participant withdraws?

Data is purged from active storage, "do not use" flags propagate to client deliverables where contractually enforceable, and the withdrawal is documented so non-use holds in future training and deliveries.

Where is the data processed and stored?

Within the EEA by default for European engagements, in the Frankfurt and Dublin regions. Transfers outside the EEA are governed by Standard Contractual Clauses where applicable.

Is YPAI the controller or the processor?

Data processor for custom collection, independent controller for licensing. The role is written into the Master Services Agreement and the DPA for each engagement.

Bring the model and the market. The consent chain comes with the corpus.

A use case, a timeline and the regulatory context are enough to start. The consent framework, the DPA and the audit package follow from scoping.