Speech data · AI Act risk classification
Four risk categories. One of them regulates the training data. Which one your system falls in, and what Article 10 then asks for.
The EU AI Act (Regulation 2024/1689) classifies AI systems into risk categories. Classification determines the obligations, including training data governance under Article 10.
Regulation 2024/1689 · Article 6 classification · Article 10 data governance
- Unacceptable Prohibited
- High Regulated. Conformity assessment, Article 10 data governance Article 10 · training, validation and testing data
- Limited Transparency obligations
- Minimal Unregulated
The four categories prohibited · high · limited · minimal
Four levels of risk. Obligations rise with the level.
The AI Act defines four levels of risk. Regulatory obligations increase with risk level.
- Unacceptable risk
- Prohibited. AI systems that pose clear threats to safety, livelihoods, or fundamental rights are banned outright: social scoring, real-time biometric identification in public spaces (with limited exceptions), AI that exploits vulnerabilities or uses subliminal manipulation. Training data is not relevant; these systems cannot be deployed.
- High risk
- Regulated. AI systems in safety-critical applications or sensitive use cases. These require conformity assessment, technical documentation, and ongoing compliance obligations including training data governance. Article 10 applies.
- Limited risk
- Transparency obligations. AI systems that interact directly with people without significant implications. Requirements focus on disclosure, so users know they are interacting with AI: chatbots, AI-generated content. Training data governance is not mandated, though good practice applies.
- Minimal risk
- Unregulated. AI systems with negligible or no impact on rights or safety, such as spam filters and AI-enabled games. No mandatory requirements.
High-risk classification article 6 · two pathways
Two pathways make a system high-risk. Profiling makes it high-risk regardless.
An AI system is classified as high-risk through two pathways defined in Article 6.
-
Pathway 1, safety component (Annex I).
The AI system is a safety component of a product, or is itself a product, covered by EU harmonization legislation listed in Annex I: medical devices, automotive systems, aviation equipment, machinery, lifts, radio equipment, and toys (where safety-relevant). If the product requires third-party conformity assessment and the AI is integral to safety, it is high-risk.
-
Pathway 2, high-risk use cases (Annex III).
The AI system falls within one of eight areas defined in Annex III.
The eight Annex III areas
- Biometrics
- Critical infrastructure
- Education and vocational training
- Employment
- Essential services
- Law enforcement
- Migration and border control
- Justice and democratic processes
- Automatic high-risk classification.
- Any AI system that profiles natural persons (automated processing of personal data to evaluate or predict aspects of a person's life, work, health, preferences, or behavior) is always classified as high-risk, regardless of exemptions.
- Exemptions from high-risk (limited).
- AI systems in Annex III areas may be exempt if they perform narrow procedural tasks, improve results of previously completed human activity, detect decision-making patterns without replacing human judgment, or perform preparatory tasks only. These exemptions do not apply if the system profiles individuals.
Article 10 quality · governance · context
What Article 10 asks of the datasets. Training, validation and testing alike.
Article 10 of the AI Act establishes data governance requirements for high-risk AI systems. These requirements apply to training, validation, and testing datasets.
Quality requirements (Article 10.3)
- Relevant to intended purpose
- Sufficiently representative
- Free of errors to best extent possible
- Complete for intended purpose
- Statistically appropriate for populations
Governance requirements (Article 10.2)
- Design choices and collection processes
- Data preparation (annotation, labeling)
- Formulation of assumptions
- Assessment of availability and suitability
- Examination for biases
- Identification of data gaps
- Measures to address identified issues
Contextual requirements (Article 10.4). Datasets must reflect the specific context of deployment: geographic setting, behavioral context, functional environment, and affected populations.
Documentation requirements. High-risk system providers must maintain technical documentation demonstrating Article 10 compliance. This documentation is subject to review during conformity assessment and may be requested by competent authorities.
Procurement sourcing decides the demonstration
The question at conformity assessment is whether training data governance can be demonstrated.
Organizations deploying high-risk AI systems must demonstrate that training data meets Article 10 requirements. Training data sourcing decisions directly affect whether that demonstration can be made.
Supports compliance
- Documented provenance and collection methodology
- Transparent sampling and representativeness information
- Bias assessment and limitations disclosure
- Version control and reproducibility
Creates compliance risk
- Lack of traceability to data sources
- No documentation of collection or preparation
- No assessment of representativeness or gaps
- No governance artifacts suitable for audit
The question during conformity assessment is not whether training data is available, but whether training data governance can be demonstrated.
YPAI's role the article 10 file
Speech datasets with the governance records an Article 10 file needs.
YPAI delivers speech and language datasets with the governance records a high-risk system's Article 10 file needs.
What YPAI provides
- European speech datasets with documented governance
- Provenance records and collection methodology
- Sampling methodology and representativeness information
- Bias assessment and known limitations disclosure
- Technical documentation for Article 10 support
How this supports Article 10. Organizations deploying high-risk AI systems can use YPAI's documentation to demonstrate training data governance during conformity assessment. The documentation is structured to address Article 10 requirements.
Whether training data meets the specific requirements for a given AI system depends on the system's intended purpose, deployment context, and affected populations.
Who decides purpose · annex · exemptions · profiling
Classification follows the system's intended purpose. Official guidance sits with the Commission.
High-risk classification depends on the system's intended purpose, whether it falls under Annex I or Annex III, whether exemptions apply, and whether the system profiles natural persons.
Official guidance sits with the European Commission and national competent authorities.
The Commission is required to publish guidelines with practical examples of high-risk and non-high-risk systems by February 2026.
Key dates 2024 · 2028
The obligations arrive in stages. Annex III systems are regulated from December 2027.
The Digital Omnibus on AI moved the high-risk dates in July 2026. Organizations deploying high-risk AI systems should have training data governance in place before 2 December 2027, and before 2 August 2028 for AI embedded in regulated products.
-
August 2024
AI Act entered into force
-
February 2025
Prohibited practices in effect
-
August 2025
GPAI model obligations in effect
-
July 2026
Digital Omnibus on AI in force, high-risk dates moved
-
August 2026
Article 50 transparency duties in effect
-
December 2027
High-risk system obligations in effect (Annex III)
-
August 2028
High-risk system obligations in effect (Annex I embedded products)
Official sources eur-lex · commission
The regulation, the service desk and the policy page. Read the text itself.
Bring the system's classification. The Article 10 records come with the corpus.
Scope speech data with the provenance, sampling and limitation records an Article 10 file needs.
Speech data overview EU AI Act compliant training data GDPR-compliant speech data DPA overview