---
title: "AI Data Validation under EU AI Act Article 10 | YPAI"
url: https://ypai.ai/data-solutions/validation/
description: "Cohen's and Fleiss' Kappa, statistical representativeness, bias detection. Every validation deliverable maps to a named EU AI Act Article 10 paragraph."
source: "src/copy/routes (route /data-solutions/validation/)"
---

# Article 10 data validation, with agreement measured and sampling documented.

> Cohen's and Fleiss' Kappa, statistical representativeness, bias detection. Every validation deliverable maps to a named EU AI Act Article 10 paragraph.

Cohen's and Fleiss' Kappa, statistical representativeness audits, bias detection. Every validation deliverable maps to a named EU AI Act paragraph.

YPAI is a Norwegian limited company (AS). Feasibility read within one EU business day.

- 2. Ten documented checks behind every validation set
- 3. Validation is evidence the conformity assessment reads
- 4. How we validate, under EU AI Act Article 10
- 6. Validation sits inside a four-service pipeline

## 1. Validation readout

- metric + threshold agreed for the task
- Art. 10(2-5), against your acceptance requirements
- a statistical sample from every delivered lot, the batch under acceptance, with gold items seeded
- two per item, adjudication on disagreement, credentialed where the domain requires it
- Cohen's Kappa for two raters, Fleiss' Kappa for three or more, Krippendorff's alpha where raters vary
- agreed per task before the lot; 0.70 and 0.85 are the engagement defaults
- ten, each with a named artefact and the paragraph it cites
- accept, remediate or replace, with every failed item named
- declaration, manifest and acceptance log ship with the lot

Each check produces a named, auditable artefact. Together they form the Article 10 governance evidence a conformity assessment expects: internal control under Annex VI for most Annex III systems, a notified body where biometrics require one.

- Demographic and contextual distribution mapping cross-referenced against contributor metadata with documented sampling methodology.
- Differential impact testing routes identical tasks through defined contributor cohorts; variances analysed and exported.
- Multi-annotator overlap quantified via Cohen's Kappa (two raters) and Fleiss' Kappa (three or more) against a pre-defined reliability threshold.
- Every validation set audited against the Article 10(3) requirement that data sets be, to the best extent possible, free of errors and complete. Error rates and gaps quantified before delivery.
- Distribution matrix maps the contributor pool against the target population across 150+ languages.
- Every label checked against the agreed task taxonomy. Out-of-schema values, ambiguous classes, inconsistent application flagged at QA.
- Validation cohorts deliberately probe rare conditions and adversarial inputs, surfacing failure modes uniform sampling leaves undocumented.
- Exact and near-duplicate detection across items and across the training, validation, and test splits, and overlap checks against the benchmarks the model is scored on. Leakage between splits is named item by item.
- Declaration maps validation outputs to EU AI Act Article 10 paragraphs 2 to 5 and Article 11, supplying the technical documentation a conformity assessment requires.
- Each delivery ships with an acceptance log and dataset manifest. Provenance, preparation, quality metrics recorded as chain-of-custody evidence.

the robot moved during exposure, so the frame cannot be labelled to the guideline

motion blur, the robot moved during exposure

## 3. Validation is evidence the conformity assessment reads.

EU AI Act Article 10 makes representative, error-free training data a statutory conformity requirement for high-risk systems. A validation gap surfaces in the conformity assessment.

Inter-rater agreement measured against the threshold agreed for the task, with a documented sampling methodology, delivered as audit-ready artefacts. Every claim maps to a named Article 10 paragraph and a named deliverable.

Four methodology stages, each mapped to a specific Article 10 paragraph. Inter-rater agreement is reported as Cohen's Kappa against a documented per-task threshold.

Agreement between reviewers, measured per task.

Two reviewers judge every sampled item without seeing each other's verdict. Disagreements go to adjudication, and the agreement between reviewers is reported as a number against the floor agreed for the task, before the lot is accepted.

### Statistical representativeness checks

### Bias detection and mitigation

### Inter-rater agreement reporting

### Article 10 conformance checkpoint

How the kappa value is read (Landis and Koch, 1977)

YPAI calibrates reliability thresholds per task. The two markers below are the documented engagement defaults; specific projects can require tighter floors.

Observed agreement between annotators on the same items.

Expected agreement by chance, given marginal class distribution.

Above the κ ≥ 0.70 high-subjectivity floor; one calibration cycle short of the κ ≥ 0.85 high-risk floor.

## 5. Every claim, mapped to a named statute

Procurement and legal teams can verify each line against the standard DPA, included with every data engagement.

Every deliverable maps to a named paragraph, and the standard DPA carries the same lines, so procurement and legal verify the page against the contract they sign.

### [Article 10](https://artificialintelligenceact.eu/article/10/)

- Training, validation, and testing datasets are assessed for relevance, representativeness, and documented errors. Bias detection and correction are documented. Human QA follows the agreed acceptance and sampling plan.

### [Article 11](https://artificialintelligenceact.eu/article/11/)

- The Data Governance Declaration details data origin, collection, and preparation. It documents how the Article 10 practices were applied during development.

### [Chapter V](https://gdpr-info.eu/chapter-5/)

- YPAI is a Norwegian limited company (AS). For EEA-pinned engagements, no third-country transfer mechanism is needed in YPAI's directly controlled processing chain; where a transfer is required, SCCs are in place. Sub-processor list and jurisdictions itemised in the DPA.
- Chap. V + DPA Art. 28

### [prEN 18284](https://www.cencenelec.eu/areas-of-work/cen-cenelec-topics/artificial-intelligence/)

- The draft European standard written to Article 10: acquisition, collection, labelling, storage, filtering, and retention of training, validation, and testing data. Not yet published or cited in the Official Journal; presumption of conformity under Article 40 follows citation. Every check above is named so it can be cross-referenced to the standard's clauses when it is cited.

## When Article 10 applies

The Digital Omnibus on AI moved the high-risk dates. Validation evidence is built before the obligation applies.

### AI Act in force.

- Regulation (EU) 2024/1689 enters into force; the high-risk chapter is written.

### Digital Omnibus on AI in force.

- Regulation (EU) 2026/1744 defers the high-risk dates and extends the Article 10(5) rule on special-category data for bias detection.

### Annex III high-risk obligations apply.

- Article 10 data governance binds providers of Annex III systems from this date. The Commission's draft classification guidelines under Article 6(5) went to consultation in May 2026.

### Annex I high-risk obligations apply.

- Systems under existing EU harmonisation legislation follow a year later.

Data validation is one component of European regulatory conformity. See how the rest of the EEA data layer composes around it.

### [Ethical Framework](https://ypai.ai/data-solutions/ethical-framework/)

- Sovereignty posture, erasure SLA, and the Norwegian-AS jurisdictional position.

### [Annotation](https://ypai.ai/data-solutions/annotation/)

- Human-in-the-loop annotation workflows with the standard DPA included by default.

### [Data Collection](https://ypai.ai/data-collection/)

- Identity-verified contributors with documented GDPR Article 7 consent and EEA processing.

### [Data Solutions hub](https://ypai.ai/data-solutions/)

- The full pipeline: collection, annotation, validation, and governance for regulated AI.

## 7. Procurement FAQ

What procurement, legal, and security ask first.

### How do you document statistical representativeness for Article 10 audits?

A demographic and contextual distribution matrix maps the human QA contributor pool against the high-risk system's intended purpose. The matrix satisfies EU AI Act Article 10(3) with documented sampling methodology.

### How do you detect bias without collecting more personal data than necessary?

Contributor metadata is used only to measure differential output variance. No extraneous personal data is processed. This aligns Article 10(2)(f) with GDPR Article 5(1)(c).

### How is inter-rater agreement calculated and reported?

Multi-annotator overlap can be quantified via Cohen's Kappa for two raters, Fleiss' Kappa for three or more, or another task-appropriate agreement method. The metric, sampling design, threshold, and achieved result are reported against the acceptance plan agreed for the task.

### Does your validation process introduce third-country data transfer risks?

YPAI is a Norwegian limited company (AS). For EEA-pinned engagements, our directly controlled processing chain introduces no third-country transfer mechanisms or Transfer Impact Assessment requirements. Sub-processor jurisdictions are itemised in the DPA so your DPO and legal team can verify the full chain of custody.

### How does your human QA map to EU AI Act documentation requirements?

Every engagement ships with a Data Governance Declaration detailing origin, collection, and preparation. This is the technical documentation required by EU AI Act Article 11.

### Are we required to establish standard contractual clauses (SCCs)?

For EEA-pinned engagements, no SCCs are needed in YPAI's directly controlled processing chain; where a customer-directed transfer requires one, SCCs are in place. Every data engagement includes a GDPR Article 28 aligned DPA, shipped with the statement of work.

### Does this validation methodology apply to LLM preference data and RLHF datasets?

Yes. Inter-rater agreement extends to RLHF and LLM evaluation: Cohen's Kappa quantifies agreement on dual-rater preference comparisons (which of two responses is preferred), and Fleiss' Kappa quantifies multi-rater consensus on output quality dimensions such as helpfulness, safety, and factuality. Representativeness checks and bias detection apply identically to preference labels and to traditional classification labels.

## Scope a validation project.

Bring the model, the operational environment, and the conformance target. We reply within one EU business day with a feasibility read, return an indicative scope, timeline, and pricing band, then deliver a Data Governance Declaration mapped to EU AI Act Article 10 paragraphs 2 to 5.

- EU AI Act Article 10 conformance, Article 11 documentation
- EEA residency by default, sub-processor list in DPA

EU AI Act Article 10 · Article 11 · GDPR Chapter V

## 9. From validation brief to a scoped pilot plan

After you submit the validation brief, we scope the metric, sample, operational environment, acceptance threshold, evidence outputs, timeline, and commercial terms.

### Project lead reads your brief.

- A named project lead replies inside one EU business day with feasibility, scope clarifications, and a first read on the Article 10 risk classification.

### Indicative scope, timeline, pricing band.

- Initial scope returned with the task-appropriate agreement method, sampling design, acceptance threshold, evidence outputs, and delivery plan.

### Scoped pilot delivered.

- The pilot uses the actual validation sample, specification, operational environment, QA method, evidence requirements, and acceptance criteria agreed for the project. Scope, timing, and commercial terms are project-specific.

### Master DPA signed, production scope locked.

- The DPA, processing locations, sub-processors, delivery plan, and production acceptance gates are agreed before scale-up.

EEA residency by default. GDPR Article 7 consent on every contributor. EU AI Act Article 10 evidence pack at delivery.

- [Data Solutions hub](https://ypai.ai/data-solutions/)
- [Annotation](https://ypai.ai/data-solutions/annotation/)
- [Data Collection](https://ypai.ai/data-collection/)
- [Ethical Framework](https://ypai.ai/data-solutions/ethical-framework/)
- [Audio data QA and acceptance criteria](https://ypai.ai/compliance/audio-data-qa/)
- [Evaluation programme](https://ypai.ai/speech-data/evaluation-program/)
