AI data annotation

Human eyes where your acceptance plan puts them.

Image, video, text, audio, LiDAR and sensor-fusion annotation under one quality regime. Marked by people, countermarked where it matters, released against your acceptance plan, delivered with the record. Governed in the EEA.

Agreement, marker A against marker B

computing IoU

Computed in this tab from the two marks on the frame.

This is an illustration. No dataset, annotator or reviewer on this page is real.

Why annotation fails in training intake · guideline · first pass · independent · agreement · reviewer · record

Labels look right until the model says otherwise.

One item, from intake to the record it is delivered with. Every value on this rail is computed from the two marks in the frame above.

  1. Intake Unopened. One frame, held in the EEA.
  2. Guideline Version 1 as written: three clauses, reference set cleared.
  3. Marker A · first pass Opened by a person. Modal: the visible pixels only.
  4. Marker B · independent Blind to A. Amodal: the body continues behind the post.
  5. Agreement Computed here. Marked a second time where the two readings differ.
  6. Reviewer Named. Boundary, major. The second reading holds.
  7. Record Delivered with the labels: guideline, agreement, DPA.
  1. The hard 20 percent hides in frames nobody opened.

    Occlusion, low light, crowds and mixed scripts hide in the frames a random sample never reaches. The agreed sampling plan names those slices, so human review reaches them on purpose.

  2. Two annotators who agree are only the start.

    Agreement is a number, scored per task against a reference set your team accepts. Where two independent markers disagree, a named reviewer decides.

  3. Labels without their record do not pass assessment.

    EU AI Act Article 10 asks for representativeness, error examination, bias examination and provenance. Each arrives as a document with the labels.

The six routes Image · Video · Segmentation · Text · Audio · LiDAR and fusion

Six kinds of raw material. One table.

Perception work leans on image, video, LiDAR and fusion. Conversational work leans on text and audio. Each modality is its own service with its own guidelines, reviewers and acceptance metric. The name opens the route.

  1. yolov10n
  2. slimsam-77
  3. electra-small-ner
  4. whisper-tiny.en
  5. depth-anything-v2
  6. Each first pass runs in your browser from our own storage.
  1. 01 Image annotation Bounding boxes, polygons and keypoints on a real scene mAP · IoU First pass · IoU Reader yolov10n · int8 · ran in this tab Market 95 %+ IoU per batch published Aug 2026 See image annotation
  2. 02 Video annotation Multi-object tracks held across frames MOTA · IDF1 · HOTA First pass · IoU, mean over 3 frames Reader yolov10n · int8 · ran in this tab Market 0.88+ IoU per frame, ID swap under 0.5 % published Aug 2026 See video annotation
  3. 03 Semantic segmentation Per-pixel masks over a real scene mIoU · PQ First pass · Dice Reader slimsam-77 · q8 · ran in this tab Market mIoU at or above 0.90 published Jul 2026 See segmentation annotation
  4. 04 Text annotation Entity spans held across mixed scripts token F1 · kappa First pass · token F1 Reader electra-small-ner · q8 · ran in this tab Market 0.92+ Cohen kappa published Aug 2026 See text annotation
  5. 05 Audio and speech Diarised speaker lanes with transcript WER · DER First pass · temporal IoU Reader whisper-tiny.en · q8 · ran in this tab Market 99.8 % accuracy on clean speech published 2026 No vendor publishes a DER ceiling. See audio annotation
  6. 06 LiDAR and sensor fusion 3D cuboids on a point cloud, projected to camera 3D IoU · cross-sensor consistency First pass · BEV IoU · relative depth Reader depth-anything-v2 · q4f16, with yolov10n · ran in this tab Market 3D IoU at or above 0.90 published Jul 2026 The detector draws a box; the open frame reads deeper than the box. Boxes and slabs disagree, which is what a second marker is for. See lidar and fusion annotation
First pass 0.825 IoU
A model's first pass against the held mark on each example, in the measure that unit uses, computed here from the geometry above. Agreement between two people is the number in the frame and the second marker.
Market 95 %+ IoU per batch published Aug 2026
The strongest level a vendor has published for that unit in 2026; sources on request. Not a YPAI commitment: your targets are set in the pilot's acceptance plan.
Reader yolov10n · int8 · in this tab
Every first pass on this table is read in your browser from our own storage, in this tab.
  1. Data you do not hold yet

    Data you do not hold yet is a collection, not an annotation. It is captured under the same quality regime and comes back onto this table.

    Scope a data collection
  2. A modality not on the table

    Working in a modality not on the table? We scope custom annotation protocols.

    Scope a project

Under the loupe Occlusion · Low light · Crowd density · Mixed scripts

Where cheap labelling lets go, the mark holds.

The hard 20 percent is where labels break. These are the failure modes our guidelines and QA are written for.

  1. Case 01

    Occlusion

    The box is held through partial visibility. It does not shrink to what is visible.

  2. Case 02

    Low light

    The mask edge follows the object, not the light.

  3. Case 03

    Crowd density

    Every instance keeps its own identity. No two figures share a box.

  4. Case 04

    Mixed scripts

    Entity spans hold across Latin, Cyrillic and Arabic in one line.

Drag the slider across a case. Left of it, the mark that held. Right of it, the model's first pass on its own. Both stand on the same crop of the frame above; the first pass is read in this tab.

Production work runs against your raw data under your engagement DPA.

The second marker first pass · independent · guideline v1 · v2

Two markers on one frame, and a number for how often they agree.

Calibration comes first The production team clears the task-specific reference set and acceptance threshold before a single production item is marked.

person · first pass person · independent
Marker A · first pass Marker B · independent marked twice where they differ

Guideline · person · box v1, as written

  1. Box every person to the visible pixels.
  2. One box per instance in a crowd; no shared boxes.
  3. A span covers the whole entity across mixed scripts.
  4. occlusion: not yet written An occluded object is labelled to its full extent. The boundary continues behind the occluder.

Agreement, marker A against marker B

0.765 IoU from 0.765

Item
Example item
Defect
boundary · major
Decision
adjudicated by a named reviewer

The number moved because the guideline did. Both are computed here from the two geometries beside it.

Disagreement goes to adjudication, not to averaging.

Agreement scored. Disagreement decided by a named reviewer.

Compliance representativeness · error-freeness · bias examination · provenance · GDPR Articles 9 and 25

EU AI Act Article 10, satisfied at the label.

Article 10 obligations cascade to the annotation partner. Each clause maps to a concrete deliverable you can hand your conformity assessor.

Governed from Norway. · Processed in the EEA. · A DPA on every engagement.

The record for the one item this page has followed, under guideline v2, written on the frame it belongs to. The evidence is the documentation: EEA jurisdiction, named regulations and audit-ready records.
  1. Article 10 representativeness Ontology and sampling design Demographic and segment distribution report
  2. Article 10 error-freeness Kappa-gated QA and gold sets Per-class agreement and defect report
  3. Article 10 bias examination Label-level bias audit Bias-examination notes per dataset
  4. Article 10 provenance Per-item provenance logging Dataset datasheet and provenance log
  5. GDPR Articles 9 and 25 Lawful basis, minimisation, data protection by design Signed DPA, 30-day erasure SLA, sub-processor list

Scale on proof, not promises Scope · Pilot · Exit gate · Production

You commit to one sheet before you commit to scale.

  1. defined with your team Scope

    Scope first: objectives, modalities, taxonomies, quality targets, risk level and DPA, defined with your team.

  2. measured, reported Pilot

    Then a measured pilot with iterative guideline refinement and reporting against the agreed metrics.

  3. targets met, process stable Exit gate

    Production only after the exit gate: metric targets met and a stable, repeatable process, with SLAs on throughput and defect ceilings.

  4. continuous QA, with SLAs Production

    Then continuous QA, relabelling campaigns and the annotation-to-model feedback loop.

The exit gate is named before scale: metric targets plus a stable process, not a single number.

Scope your pilot Image · Video · Segmentation · Text · Audio · LiDAR and fusion

Put your data on the table.

We scope the modalities, taxonomies and quality targets your deployment needs. As a managed European partner, YPAI carries the QA, project management and compliance work.

What the brief needs: modality, volume, taxonomy, and what acceptance looks like.

Send the brief An engineer or delivery lead replies inside one EU business day with a feasibility read and the next concrete step.

Governed from Norway. · Processed in the EEA. · A DPA on every engagement.

Read the data residency brief →

FAQ

Frequently asked questions

What types of AI data annotation does YPAI support?

We annotate images (bounding boxes, polygons, semantic segmentation, keypoints), video (object tracking, event labelling), text (named-entity recognition, intent, sentiment), audio and speech (transcription, diarisation, prosody), and 3D / LiDAR point clouds for autonomous-vehicle perception. Every modality runs under one quality and compliance regime, so a multimodal project shares one review, provenance and audit-trail pipeline. AI data labelling covers how labelling and annotation differ.

How does YPAI ensure annotation quality?

Each project runs spec calibration, a second pass on dispute-prone slices, expert adjudication, and a final QA sweep against the project's acceptance criteria. Every batch ships with an inter-annotator agreement (IAA) report and a per-label confusion matrix, so your team re-works the labels that move training. Sample sizes are scoped statistically to the project.

Is YPAI annotation GDPR-compliant?

Annotation pipelines are built under GDPR from the first scoping call: lawful basis, purpose limitation, data minimisation and Article 35 DPIA scaffolding are part of the engagement. Personal data is processed on EU/EEA infrastructure under a signed Data Processing Agreement (DPA), and subject-rights requests (access, rectification, erasure) go through the data request form. The data ethical framework sets out the governance behind it.

Which languages does YPAI cover for text and speech annotation?

YPAI supports annotation across 150+ languages and dialects, with native-speaker reviewers concentrated in European, Nordic and major Asian markets. Lower-resource languages are quoted per project, because reviewer recruitment sets the pace there. A multilingual project runs through one delivery lead, so glossary, style guide and IAA targets hold across languages.

Can YPAI annotate medical imaging data (DICOM, FHIR)?

We handle DICOM, NIfTI and FHIR-bundled imaging with clinician reviewers contracted for the modality. Data is processed inside EEA residency by default, and the pipeline supports GDPR Article 9 special-category handling. The statement of work documents the exact controls and evidence artefacts.

What is the minimum project size YPAI accepts?

We take on work where the annotation is non-trivial and the compliance posture matters. The smallest engagement that usually justifies scoping is a single-batch pilot: a few thousand items for image or text, a small corpus for speech. Pilots are common, and the scoping call confirms fit before any quote.

How fast does YPAI reply after a project enquiry?

We reply inside one EU business day after a submission to /contact-us/, with a feasibility read, the next concrete step and an estimated scope window. The first reply comes from an engineer or delivery lead who has read the brief.

How is YPAI different from Scale AI, Labelbox, or Appen?

YPAI is headquartered in Lysaker, Norway, and operates inside the EU/EEA, as a Norwegian company with no US corporate entity. Data residency, subprocessors and transfer controls are defined per project. In-house teams with sector specialists (automotive OEM, healthcare ASR, financial documents) run each engagement. Pricing is set per project after a feasibility scope.

Human eyes where your acceptance plan puts them.

Scope your annotation pilot

The form adapts to the work, asks only for relevant details and sends your brief to the person who can act on it.

Service required

Your selection routes the brief to the right person.

A named project lead reviews every enquiry and replies within one business day