---
title: "AI Data Annotation Across Every Modality | YPAI"
url: https://ypai.ai/ai-data-annotation/
description: "YPAI labels image, video, text, audio, LiDAR, and sensor-fusion data with kappa-gated QA and human review, governed in the EEA under EU AI Act Article 10."
source: "src/copy/routes (route /ai-data-annotation/)"
---

# Human eyes where your acceptance plan puts them.

> YPAI labels image, video, text, audio, LiDAR, and sensor-fusion data with kappa-gated QA and human review, governed in the EEA under EU AI Act Article 10.

Image, video, text, audio, LiDAR and sensor-fusion annotation under one quality regime. Marked by people, countermarked where it matters, released against your acceptance plan, delivered with the record. Governed in the EEA.

## Labels look right until the model says otherwise.

### The hard 20 percent hides in frames nobody opened.

- Occlusion, low light, crowds and mixed scripts hide in the frames a random sample never reaches. The agreed sampling plan names those slices, so human review reaches them on purpose.

### Two annotators who agree are only the start.

- Agreement is a number, scored per task against a reference set your team accepts. Where two independent markers disagree, a named reviewer decides.

### Labels without their record do not pass assessment.

- EU AI Act Article 10 asks for representativeness, error examination, bias examination and provenance. Each arrives as a document with the labels.

## Six kinds of raw material. One table.

Perception work leans on image, video, LiDAR and fusion. Conversational work leans on text and audio. Each modality is its own service with its own guidelines, reviewers and acceptance metric. The name opens the route.

Data you do not hold yet is a collection, not an annotation. It is captured under the same quality regime and comes back onto this table.

Working in a modality not on the table? We scope custom annotation protocols.

- Bounding boxes, polygons and keypoints on a real scene
- Per-pixel masks over a real scene
- Entity spans held across mixed scripts
- 3D cuboids on a point cloud, projected to camera

## Where cheap labelling lets go, the mark holds.

The hard 20 percent is where labels break. These are the failure modes our guidelines and QA are written for.

- The box is held through partial visibility. It does not shrink to what is visible.
- The mask edge follows the object, not the light.
- Every instance keeps its own identity. No two figures share a box.
- Entity spans hold across Latin, Cyrillic and Arabic in one line.

Production work runs against your raw data under your engagement DPA.

## Two markers on one frame, and a number for how often they agree.

- Dispute-prone slices are annotated twice, independently. Agreement is scored with the method chosen for the task, against a reference set your team accepts. Disagreement goes to adjudication, not to averaging.
- Every defect is classified by type (boundary, class, attribute, miss) and severity, and reported per class as a confusion matrix, so the next guideline revision has somewhere precise to begin.

The production team clears the task-specific reference set and acceptance threshold before a single production item is marked.

Agreement scored. Disagreement decided by a named reviewer.

## EU AI Act Article 10, satisfied at the label.

Article 10 obligations cascade to the annotation partner. Each clause maps to a concrete deliverable you can hand your conformity assessor.

The evidence is the documentation: EEA jurisdiction, named regulations and audit-ready records.

- Lawful basis, minimisation, data protection by design
- Signed DPA, 30-day erasure SLA, sub-processor list

## You commit to one sheet before you commit to scale.

Scope first: objectives, modalities, taxonomies, quality targets, risk level and DPA, defined with your team. Then a measured pilot with iterative guideline refinement and reporting against the agreed metrics. Production only after the exit gate: metric targets met and a stable, repeatable process, with SLAs on throughput and defect ceilings. Then continuous QA, relabelling campaigns and the annotation-to-model feedback loop.

The exit gate is named before scale: metric targets plus a stable process, not a single number.

## Put your data on the table.

We scope the modalities, taxonomies and quality targets your deployment needs. As a managed European partner, YPAI carries the QA, project management and compliance work.

## What types of AI data annotation does YPAI support?

We annotate images (bounding boxes, polygons, semantic segmentation, keypoints), video (object tracking, event labelling), text (named-entity recognition, intent, sentiment), audio and speech (transcription, diarisation, prosody), and 3D / LiDAR point clouds for autonomous-vehicle perception. Every modality runs under one quality and compliance regime, so a multimodal project shares one review, provenance and audit-trail pipeline. <a href="/ai-data-labeling/">AI data labelling</a> covers how labelling and annotation differ.

## How does YPAI ensure annotation quality?

Each project runs spec calibration, a second pass on dispute-prone slices, expert adjudication, and a final QA sweep against the project's acceptance criteria. Every batch ships with an inter-annotator agreement (IAA) report and a per-label confusion matrix, so your team re-works the labels that move training. Sample sizes are scoped statistically to the project.

## Is YPAI annotation GDPR-compliant?

Annotation pipelines are built under GDPR from the first scoping call: lawful basis, purpose limitation, data minimisation and Article 35 DPIA scaffolding are part of the engagement. Personal data is processed on EU/EEA infrastructure under a signed Data Processing Agreement (DPA), and subject-rights requests (access, rectification, erasure) go through the <a href="/gdpr/request/">data request form</a>. The <a href="/data-solutions/ethical-framework/">data ethical framework</a> sets out the governance behind it.

## Which languages does YPAI cover for text and speech annotation?

YPAI supports annotation across 150+ languages and dialects, with native-speaker reviewers concentrated in European, Nordic and major Asian markets. Lower-resource languages are quoted per project, because reviewer recruitment sets the pace there. A multilingual project runs through one delivery lead, so glossary, style guide and IAA targets hold across languages.

## Can YPAI annotate medical imaging data (DICOM, FHIR)?

We handle DICOM, NIfTI and FHIR-bundled imaging with clinician reviewers contracted for the modality. Data is processed inside EEA residency by default, and the pipeline supports GDPR Article 9 special-category handling. The statement of work documents the exact controls and evidence artefacts.

## What is the minimum project size YPAI accepts?

We take on work where the annotation is non-trivial and the compliance posture matters. The smallest engagement that usually justifies scoping is a single-batch pilot: a few thousand items for image or text, a small corpus for speech. Pilots are common, and the scoping call confirms fit before any quote.

## How fast does YPAI reply after a project enquiry?

We reply inside one EU business day after a submission to <a href="/contact-us/">/contact-us/</a>, with a feasibility read, the next concrete step and an estimated scope window. The first reply comes from an engineer or delivery lead who has read the brief.

## How is YPAI different from Scale AI, Labelbox, or Appen?

YPAI is headquartered in Lysaker, Norway, and operates inside the EU/EEA, as a Norwegian company with no US corporate entity. Data residency, subprocessors and transfer controls are defined per project. In-house teams with sector specialists (automotive OEM, healthcare ASR, financial documents) run each engagement. Pricing is set per project after a feasibility scope.
