AI DATA & EVALUATION Validation conformitycoverageduplicationrights

Article 10 data validation, with agreement measured and sampling documented.

scope
Art. 10(2-5), against your acceptance requirements
sample
a statistical sample from every delivered lot, the batch under acceptance, with gold items seeded
reviewers
two per item, adjudication on disagreement, credentialed where the domain requires it
agreement
Cohen's Kappa for two raters, Fleiss' Kappa for three or more, Krippendorff's alpha where raters vary
threshold
agreed per task before the lot; 0.70 and 0.85 are the engagement defaults
checks
ten, each with a named artefact and the paragraph it cites
decision
accept, remediate or replace, with every failed item named
record
declaration, manifest and acceptance log ship with the lot
Specialised validation
Audio data QASpeech specificationsProvenance auditDataset auditsLabel QA and agreement analysisRepresentativeness and bias reviewDuplication and contamination checksConsent and rights verificationAcceptance sampling and gold setsSensor calibration and alignmentSynthetic-to-real validationVideo integrity and synthetic mediaArticle 10 data governanceEEA data residency

Inter-rater agreement measured against the threshold agreed for the task, with a documented sampling methodology, delivered as audit-ready artefacts.

You cannot trust the data you have.

the inspection sampled inspected adjudicated decided 3 to remediation

lot
sampled lot · 320 items · 16 inspected
A8 · conformity
motion blur, the robot moved during exposure
B2 · duplication
the same frame delivered twice
B8 · rights
contributor consent not on file
decision
3 items → remediation · the lot is re-inspected before acceptance · accept, remediate or replace

every failed item named, with its reason

Validation is evidence the conformity assessment reads.

Each check produces a named, auditable artefact. Together they form the Article 10 governance evidence a conformity assessment expects: internal control under Annex VI for most Annex III systems, a notified body where biometrics require one.

Ten documented checks behind every validation set
  1. Statistical representativeness Art. 10(3) representativeness_matrix.pdf
  2. Bias detection and mitigation Art. 10(2)(f) bias_mitigation_report.pdf
  3. Inter-rater agreement Art. 10(3) iaa_report.pdf
  4. Error-rate and completeness audit Art. 10(3) completeness_audit.csv
  5. Demographic and functional distribution Art. 10(3) distribution_matrix.pdf
  6. Label-taxonomy conformance Art. 10(3) taxonomy_conformance.csv
  7. Edge-case and adversarial review Art. 10(2)(g) edgecase_log.csv
  8. Duplication and contamination Art. 10(3) duplication_report.csv
  9. Data Governance Declaration Art. 10(2-5) + Art. 11 data_governance_declaration.pdf
  10. Acceptance log and manifest Art. 11 manifest.json + acceptance_log.csv

representativeness_matrix.pdf

Statistical representativeness

Demographic and contextual distribution mapping cross-referenced against contributor metadata with documented sampling methodology.

cohortitemsconformitycoverageduplicationrights
row A87888
row B88877

Article cited Art. 10(3)

bias_mitigation_report.pdf

Bias detection and mitigation

Differential impact testing routes identical tasks through defined contributor cohorts; variances analysed and exported.

cohortitemsfailedrate
row A810.13
row B820.25
difference10.13

Article cited Art. 10(2)(f)

iaa_report.pdf

Inter-rater agreement

Multi-annotator overlap quantified via Cohen's Kappa (two raters) and Fleiss' Kappa (three or more) against a pre-defined reliability threshold.

items
25
agree
23
observed
0.92
expected
0.50
kappa
0.84
floor
κ ≥ 0.70

Article cited Art. 10(3)

completeness_audit.csv

Error-rate and completeness audit

Every validation set audited against the Article 10(3) requirement that data sets be, to the best extent possible, free of errors and complete. Error rates and gaps quantified before delivery.

  • A1
  • A2
  • A3
  • A4
  • A5
  • A6
  • A7
  • A8 conformity
  • B1
  • B2 duplication
  • B3
  • B4
  • B5
  • B6
  • B7
  • B8 rights

Article cited Art. 10(3)

distribution_matrix.pdf

Demographic and functional distribution

Distribution matrix maps the contributor pool against the target population across 150+ languages.

frame quadrantitems sampled
upper left0
upper right1
lower left9
lower right6

Article cited Art. 10(3)

taxonomy_conformance.csv

Label-taxonomy conformance

Every label checked against the agreed task taxonomy. Out-of-schema values, ambiguous classes, inconsistent application flagged at QA.

classes
4
labels checked
16
out of schema
0
ambiguous
0
flagged at QA
3

Article cited Art. 10(3)

edgecase_log.csv

Edge-case and adversarial review

Validation cohorts deliberately probe rare conditions and adversarial inputs, surfacing failure modes uniform sampling leaves undocumented.

  • A1 low light
  • A5 low light
  • B1 low light
  • B7 low light

Article cited Art. 10(2)(g)

duplication_report.csv

Duplication and contamination

Exact and near-duplicate detection across items and across the training, validation, and test splits, and overlap checks against the benchmarks the model is scored on. Leakage between splits is named item by item.

itemduplicate ofmethod
B2A2exact match

Article cited Art. 10(3)

data_governance_declaration.pdf

Data Governance Declaration

Declaration maps validation outputs to EU AI Act Article 10 paragraphs 2 to 5 and Article 11, supplying the technical documentation a conformity assessment requires.

paragraphartefact
Art. 10(3)representativeness_matrix.pdf
Art. 10(2)(f)bias_mitigation_report.pdf
Art. 10(3)iaa_report.pdf
Art. 10(2-5) + Art. 11data_governance_declaration.pdf
Art. 11manifest.json + acceptance_log.csv

Article cited Art. 10(2-5) + Art. 11

manifest.json + acceptance_log.csv

Acceptance log and manifest

Each delivery ships with an acceptance log and dataset manifest. Provenance, preparation, quality metrics recorded as chain-of-custody evidence.

  • A1
  • A2
  • A3
  • A4
  • A5
  • A6
  • A7
  • A8 conformity
  • B1
  • B2 duplication
  • B3
  • B4
  • B5
  • B6
  • B7
  • B8 rights

Article cited Art. 11

Agreement between reviewers, measured per task.

Two reviewers judge every sampled item without seeing each other's verdict. Disagreements go to adjudication, and the agreement between reviewers is reported as a number against the floor agreed for the task, before the lot is accepted.

Worked example, dual-annotator review click a verdict in reviewer B's row to change it
reviewer A
reviewer B
agree
po
0.92 92% of items agreed Observed agreement between annotators on the same items.
pe
0.50 balanced binary baseline Expected agreement by chance, given marginal class distribution.
κ
0.84 Almost perfect agreement (po − pe) / (1 − pe)
How the kappa value is read (Landis and Koch, 1977)
Slight 0.00 to 0.20 Fair 0.21 to 0.40 Moderate 0.41 to 0.60 Substantial 0.61 to 0.80 Almost perfect 0.81 to 1.00 κ ≥ 0.70 High-subjectivity annotation κ ≥ 0.85 High-risk classification 0.84

YPAI calibrates reliability thresholds per task. The two markers below are the documented engagement defaults; specific projects can require tighter floors.

Above the κ ≥ 0.70 high-subjectivity floor; one calibration cycle short of the κ ≥ 0.85 high-risk floor.

Every claim, mapped to a named statute

Procurement and legal teams can verify each line against the standard DPA, included with every data engagement.

EU AI ACT Article 10 Regulation (EU) 2024/1689 read the text
Scope Data and data governance
What YPAI ships Training, validation, and testing datasets are assessed for relevance, representativeness, and documented errors. Bias detection and correction are documented. Human QA follows the agreed acceptance and sampling plan.
Deliverable Data Governance Declaration § 10(2-5)
EU AI ACT Article 11 Regulation (EU) 2024/1689 read the text
Scope Technical documentation
What YPAI ships The Data Governance Declaration details data origin, collection, and preparation. It documents how the Article 10 practices were applied during development.
Deliverable Technical documentation file § 11 + Annex IV
GDPR Chapter V Regulation (EU) 2016/679 read the text
Scope Third-country transfer
What YPAI ships YPAI is a Norwegian limited company (AS). For EEA-pinned engagements, no third-country transfer mechanism is needed in YPAI's directly controlled processing chain; where a transfer is required, SCCs are in place. Sub-processor list and jurisdictions itemised in the DPA.
Deliverable Sub-processor list in DPA Chap. V + DPA Art. 28
CEN-CENELEC prEN 18284 JTC 21 draft standard read the text
Scope Quality and governance of datasets
What YPAI ships The draft European standard written to Article 10: acquisition, collection, labelling, storage, filtering, and retention of training, validation, and testing data. Not yet published or cited in the Official Journal; presumption of conformity under Article 40 follows citation. Every check above is named so it can be cross-referenced to the standard's clauses when it is cited.
Deliverable Clause cross-reference on citation Art. 40 on citation

Every deliverable maps to a named paragraph, and the standard DPA carries the same lines, so procurement and legal verify the page against the contract they sign.

When Article 10 applies

The Digital Omnibus on AI moved the high-risk dates. Validation evidence is built before the obligation applies.

1 August 2024
AI Act in force. Regulation (EU) 2024/1689 enters into force; the high-risk chapter is written.
27 July 2026
Digital Omnibus on AI in force. Regulation (EU) 2026/1744 defers the high-risk dates and extends the Article 10(5) rule on special-category data for bias detection.
2 December 2027
Annex III high-risk obligations apply. Article 10 data governance binds providers of Annex III systems from this date. The Commission's draft classification guidelines under Article 6(5) went to consultation in May 2026.
2 August 2028
Annex I high-risk obligations apply. Systems under existing EU harmonisation legislation follow a year later.

From validation brief to a scoped pilot plan

After you submit the validation brief, we scope the metric, sample, operational environment, acceptance threshold, evidence outputs, timeline, and commercial terms.

Inside one EU business day.
Project lead reads your brief. A named project lead replies inside one EU business day with feasibility, scope clarifications, and a first read on the Article 10 risk classification.
During scoping.
Indicative scope, timeline, pricing band. Initial scope returned with the task-appropriate agreement method, sampling design, acceptance threshold, evidence outputs, and delivery plan.
After scoping.
Scoped pilot delivered. The pilot uses the actual validation sample, specification, operational environment, QA method, evidence requirements, and acceptance criteria agreed for the project. Scope, timing, and commercial terms are project-specific.
By agreement.
Master DPA signed, production scope locked. The DPA, processing locations, sub-processors, delivery plan, and production acceptance gates are agreed before scale-up.

EEA residency by default. GDPR Article 7 consent on every contributor. EU AI Act Article 10 evidence pack at delivery.

Scope a validation project.

Bring the model, the operational environment, and the conformance target. We reply within one EU business day with a feasibility read, return an indicative scope, timeline, and pricing band, then deliver a Data Governance Declaration mapped to EU AI Act Article 10 paragraphs 2 to 5.

EU AI Act Article 10 conformance, Article 11 documentation
Cohen's and Fleiss' Kappa reporting
Representativeness, bias, error-rate, distribution checks
EEA residency by default, sub-processor list in DPA

EU AI Act Article 10 · Article 11 · GDPR Chapter V

We use the information you provide to assess, respond to and manage your enquiry in accordance with our Privacy Policy.

FAQ

Procurement FAQ

How do you document statistical representativeness for Article 10 audits?

A demographic and contextual distribution matrix maps the human QA contributor pool against the high-risk system's intended purpose. The matrix satisfies EU AI Act Article 10(3) with documented sampling methodology.

How do you detect bias without collecting more personal data than necessary?

Contributor metadata is used only to measure differential output variance. No extraneous personal data is processed. This aligns Article 10(2)(f) with GDPR Article 5(1)(c).

How is inter-rater agreement calculated and reported?

Multi-annotator overlap can be quantified via Cohen's Kappa for two raters, Fleiss' Kappa for three or more, or another task-appropriate agreement method. The metric, sampling design, threshold, and achieved result are reported against the acceptance plan agreed for the task.

Does your validation process introduce third-country data transfer risks?

YPAI is a Norwegian limited company (AS). For EEA-pinned engagements, our directly controlled processing chain introduces no third-country transfer mechanisms or Transfer Impact Assessment requirements. Sub-processor jurisdictions are itemised in the DPA so your DPO and legal team can verify the full chain of custody.

How does your human QA map to EU AI Act documentation requirements?

Every engagement ships with a Data Governance Declaration detailing origin, collection, and preparation. This is the technical documentation required by EU AI Act Article 11.

Are we required to establish standard contractual clauses (SCCs)?

For EEA-pinned engagements, no SCCs are needed in YPAI's directly controlled processing chain; where a customer-directed transfer requires one, SCCs are in place. Every data engagement includes a GDPR Article 28 aligned DPA, shipped with the statement of work.

Does this validation methodology apply to LLM preference data and RLHF datasets?

Yes. Inter-rater agreement extends to RLHF and LLM evaluation: Cohen's Kappa quantifies agreement on dual-rater preference comparisons (which of two responses is preferred), and Fleiss' Kappa quantifies multi-rater consensus on output quality dimensions such as helpfulness, safety, and factuality. Representativeness checks and bias detection apply identically to preference labels and to traditional classification labels.