AI data and evaluation · Data labelling

A label is a decision, with a name on it.

Data labelling services for text, image, audio and video. Label noise starts where a definition lets two people read one item two ways. We write the definitions with you, put two people on every item, settle the splits in writing and deliver the record with the labels.

Written definitions · Two annotators, always · Tie-break on overlap · EEA delivery under DPA

First pass, a model

0.80 truck

The model is sure. The list offers car and truck, and this is a van.

Example frame, generated for this page. The detector runs in this tab; the three answers are the ones people give.

Labelling and annotation where · what

Annotation marks where it is. Labelling decides what it is.

Both words get used for both jobs, and we deliver both. A mark fails on position: the edge, the occlusion, the dark frame. A class fails on judgement: two definitions that overlap, or an item with no class on the list.

Drag the slider. Left of it is annotation, where the vehicle is. Right of it is labelling, what the vehicle is.

The detector splits its score between two classes it knows. This is a cargo bike, and the list has no cargo bike. That is a gap in the taxonomy. It is closed in writing, with a new class or a rule for where cargo bikes go, before the batch runs. Read about AI data annotation →

The definition decides three clauses · v1 · v2 · one rule

Two people, one item, two answers. The definition has a hole.

One customer message carries three asks, and each ask fits a different class. Version 1 of the definitions leaves the order open, so the label depends on who opens the item. Version 2 adds one sentence that settles it.

Example message I was charged again billing dispute 0.49 after I cancelled last month cancel subscription 0.31 and I want that money back. refund request 0.54

Classes that fit this message

from 3 3 classes

three asks, three defensible labels · margin, first class to second 0.08

clause 1 clause 2 clause 3 Each clause scored against definitions v1 · all-MiniLM-L6-v2 · precomputed

Definitions · customer enquiries v1, as written

  1. Refund request. The writer wants money back for something already paid. Use it whenever returning money is the ask.
  2. Cancel subscription. The writer wants to end a subscription or stop being billed. If the message asks for both, it is a refund request.
  3. Billing dispute. The writer says a charge is wrong. If the charge is accepted and the money is wanted back, it is a refund request.
  4. More than one ask: left to whoever opens it. More than one ask: label what we must act on first.

The full taxonomy in this example has seven classes.

A tie-break is a sentence. It is written once and it decides every item after it.

Under v2 a person applies the rule and the message is a refund request. The embedding model still ranks billing dispute first, and that disagreement goes into the log with the label.

Example taxonomy and example messages, written for this page.

Before anyone labels seven definitions · three closest pairs · cosine

We find the classes that will collide before the first item is labelled.

An embedding model reads the seven definitions and measures how close each pair sits. The closest pairs are the candidates for a sharper definition, and those get rewritten first. Whether people actually split on them is measured next, in the calibration batch.

The closest pair

0.63 cosine

refund request · return or exchange · all-MiniLM-L6-v2 · precomputed

  1. Refund request · Return or exchange 0.63
  2. Refund request · Cancel subscription 0.57
  3. Return or exchange · Damage claim 0.53

21 pairs · lowest 0.10 · median 0.38 · highest 0.63

The model finds the weak definitions. Two people label every item. A third settles the splits.

The batch, before labelling sixty messages · seven classes · the rare class

Sixty messages on the belt. The rare class is pulled out first.

The first sixty enquiries of an example batch, embedded in this tab one by one. Before labelling starts, the items nearest each thin class are pulled forward, so the class that matters most has enough labelled examples to measure agreement on.

Nearest the damage definition

8 of sixty

A class at a few percent leaves a random sample with a handful of items, too few to score. Pulling the nearest candidates first gives it enough.

  1. inq_043 The seal was broken and the contents had spilled.
  2. inq_003 The box arrived crushed and the lid is cracked right through.
  3. inq_011 The screen was shattered when I opened the package.
  4. inq_031 The frame arrived bent and the glass is loose inside.
  5. inq_038 Water got in and the box was soaked when it reached me.
  6. inq_049 Screen cracked in transit. Photos attached.
  7. inq_057 Please refund the order that never shipped.
  8. inq_022 How do I return an item that does not fit?

The small model stands in for yours. In a programme the same pull runs on your model's embeddings.

Label types image class · condition grade · intent · policy · preference

Five kinds of label, each with the agreement statistic that fits it.

Percent agreement looks high on any task, because agreement by chance is inside it. Each label type is reported with a chance-corrected statistic and the published band to read it against.

Each first pass runs in your browser from our own storage.

  1. 01

    Object class on an image

    One vehicle, one box, two classes on it.

    0.73 car · detector score

    Reader yolov10n · int8

    Level Cohen's kappa · 0.61 to 0.80 substantial · Landis and Koch

    Landis and Koch bands for kappa · Krippendorff's thresholds for alpha

  2. 02

    Condition grade on an image

    One box, graded 0 to 3.

    4 grades, 0 to 3 · three evidence pins

    Reader people only

    Level weighted kappa · 2 against 3 counts as a near miss, 0 against 3 as a full miss

  3. 03

    Intent on a message

    Parcel says delivered but nothing is at the door.

    0.44 delivery status · next damage claim 0.34

    Reader all-MiniLM-L6-v2 · q8

    Level Cohen's kappa, per class · reported per class, the rare class first

  4. 04

    Policy labels on a review

    The driver was rude and the box was soaking wet.

    0.46 courier conduct · and item damaged 0.45

    Reader all-MiniLM-L6-v2 · q8

    Level Krippendorff's alpha, set-valued · 0.80 reliable · 0.67 tentative · Krippendorff

  5. 05

    Preference between two replies

    Two replies to inq_002. Which one goes out?

    3 rubric clauses · one of them decides

    Reader people only

    Level agreement and kappa on pairs · published labeler agreement on preference data sits near three in four

The programme Taxonomy · Definitions v1 · Calibration batch · Double labelling · Adjudication · Gold · Delivery

Volume starts after three people have read the definitions the same way.

Seven stages and one gate. In the calibration batch three people label the same items independently, their agreement is measured per class, and production volume is committed once the rare class clears it.

  1. Taxonomy The classes you have, or the ones still being argued about.
  2. Definitions v1 One sentence per class, written before labelling starts.
  3. Calibration batch Three people label the same items independently.
  4. definitions held · people agreed Double labelling Two people on every production item.
  5. Adjudication Every split goes to a reviewer, with the rule applied and the reviewer's id.
  6. Gold Items with a known answer, seeded through the batch and scored per annotator.
  7. Delivery The labels and their record, versioned.

The gate · agreement on the rare class, per version, before production

What you receive Label set · Definition changes · Agreement · Adjudication log · Residency

The record is the product.

EU AI Act Article 10 asks for representativeness, error examination, bias examination and provenance. The definitions in force, the agreement per class and the split log arrive as files, so your assessor reads the history with the labels.

Label record
Example message
Definitions
v2 · one rule added
Rule
line 4 · what we must act on first
Decision
refund request · a person
Readings
three clauses · three classes
Residency
EEA · DPA on the engagement
The record for the example message under definitions v2, written on the message it belongs to. Votes and agreement per class are added when the calibration round has run.
  1. Label set Every class with the definition in force The definition file, versioned, with the tie-break.
  2. Definition changes What changed, when, and what it fixed The diff between versions and the round that motivated it.
  3. Agreement Per class, with gold seeded through the batch Pairwise and pooled kappa, alpha, the rare class first.
  4. Adjudication log Every split, decided in writing The decision, the rule, the reviewer's id, and every vote kept.
  5. Residency EEA under a signed DPA The processing record and the sub-processor list.

Scope your labelling programme class set · volume · languages · acceptance

Bring the taxonomy you have, or the one you are still arguing about.

We scope the class set, the volume, the languages and the acceptance plan. When the definitions are still open, writing them is the first work package, and the calibration batch tests them before volume.

What the brief needs: class set, volume, languages, and what acceptance looks like.

Send the brief

An engineer or delivery lead replies inside one EU business day with a feasibility read.

FAQ

Frequently asked questions

What is the difference between data labelling and data annotation?

Both words are used for both jobs across the market, and we deliver both. What differs is how they fail. A mark fails on position: the edge, the occlusion, the dark frame. A class fails on judgement: an ambiguous definition, two categories that overlap, a taxonomy that drifts between people, a rare class nobody sees. Labelling programmes are scoped on this route, and spatial marking on AI data annotation.

How does YPAI decide a case that sits between two classes?

The guideline carries a tie-break rule that says which class wins when more than one fits. When two classes still fit after the rule, the item goes to a named reviewer, and the decision, the rule applied and the reviewer's id are written into the adjudication log that ships with the batch.

How many people label each item?

Two, always, with a third who settles disagreement. Agreement between the two is measured and reported per class. In the calibration batch three people label the same items, so every split has a majority and a minority to read.

Which languages does YPAI cover?

YPAI supports work across 150+ languages and dialects, with native-speaker reviewers concentrated in European, Nordic and major Asian markets. Lower-resource languages are quoted per project, because reviewer recruitment sets the pace there.

Do you use models to pre-label?

Where it helps, as model-assisted annotation with human verification. A model's output is a first pass, and a person accepts or corrects every item before delivery. The models on this page do a different job: they read definitions and messages to show where a decision is weak.

Where is the data processed?

Inside EU/EEA infrastructure under a signed Data Processing Agreement, with lawful basis, purpose limitation and data minimisation set at scoping. Subject-rights requests route through the data request form.

How fast does YPAI reply after an enquiry?

We reply inside one EU business day after a submission through the contact form, with a feasibility read, the next concrete step and an estimated scope window. The first reply comes from an engineer or delivery lead who has read the brief.

A label is a decision. A decision needs a definition, a tie-break and a record of who made it.

Scope your labelling programme

The form adapts to the work, asks only for relevant details and sends your brief to the person who can act on it.

Service required

Your selection routes the brief to the right person.

A named project lead reviews every enquiry and replies within one business day