---
title: "Enterprise AI Data Labeling Services | YPAI"
url: https://ypai.ai/ai-data-labeling/
description: "Data labeling with written class definitions, a tie-break for overlapping classes, two annotators on every item and the record of who decided. EEA, under DPA."
source: "src/copy/routes (route /ai-data-labeling/)"
---

# A label is a decision, with a name on it.

> Data labeling with written class definitions, a tie-break for overlapping classes, two annotators on every item and the record of who decided. EEA, under DPA.

Data labelling services for text, image, audio and video. Label noise starts where a definition lets two people read one item two ways. We write the definitions with you, put two people on every item, settle the splits in writing and deliver the record with the labels.

- van · not on the list

Example item · opened by a person · processed in the EEA

The model is sure. The list offers car and truck, and this is a van.

Example frame, generated for this page. The detector runs in this tab; the three answers are the ones people give.

## Annotation marks where it is. Labelling decides what it is.

Both words get used for both jobs, and we deliver both. A mark fails on position: the edge, the occlusion, the dark frame. A class fails on judgement: two definitions that overlap, or an item with no class on the list.

the class · first pass only

The detector splits its score between two classes it knows. This is a cargo bike, and the list has no cargo bike. That is a gap in the taxonomy. It is closed in writing, with a new class or a rule for where cargo bikes go, before the batch runs.

Drag the slider. Left of it is annotation, where the vehicle is. Right of it is labelling, what the vehicle is.

## Two people, one item, two answers. The definition has a hole.

three clauses · v1 · v2 · one rule

One customer message carries three asks, and each ask fits a different class. Version 1 of the definitions leaves the order open, so the label depends on who opens the item. Version 2 adds one sentence that settles it.

More than one ask: left to whoever opens it.

Example taxonomy and example messages, written for this page.

Each clause scored against definitions v1

Under v2 a person applies the rule and the message is a refund request. The embedding model still ranks billing dispute first, and that disagreement goes into the log with the label.

one rule, applied by a person

More than one ask: label what we must act on first.

The full taxonomy in this example has seven classes.

A tie-break is a sentence. It is written once and it decides every item after it.

- The writer wants money back for something already paid.
- Use it whenever returning money is the ask.
- The writer asks for money already paid to be returned. Use this class whenever returning money is the ask, even if the writer also wants the subscription stopped.
- The writer wants to end a subscription or stop being billed.
- If the message asks for both, it is a refund request.
- The writer asks to stop future billing and asks for no money back. If the message asks for both, it is a refund request.
- The writer says a charge is wrong.
- If the charge is accepted and the money is wanted back, it is a refund request.
- The writer says a charge is wrong and asks us to check it: the amount, the date, the count, or a charge they do not recognise. If the writer accepts that the charge happened and wants the money returned, it is a refund request.
- The writer asks where an order is or when it will arrive.
- The writer wants to send an item back or swap it for another.
- The writer wants to send an item back or swap it, and the item is intact. An item that arrived broken is a damage claim.
- The writer cannot get into their account.
- The writer says an item arrived damaged.
- The writer says an item arrived damaged, broken, leaking or otherwise unusable on arrival. This class holds even when the writer also asks for money back or a replacement, because the damage is what we must act on.

I was charged again after I cancelled last month and I want that money back.

- and I want that money back.

When a message carries more than one ask, label it by what we must act on first: damage before return, money back before stopping billing, and checking a charge before returning it. If two classes still fit, it goes to adjudication and the adjudicator's reason is recorded.

## We find the classes that will collide before the first item is labelled.

seven definitions · three closest pairs · cosine

An embedding model reads the seven definitions and measures how close each pair sits. The closest pairs are the candidates for a sharper definition, and those get rewritten first. Whether people actually split on them is measured next, in the calibration batch.

The model finds the weak definitions. Two people label every item. A third settles the splits.

Under v2 the pairs sit closer, because each definition now names its neighbour. Closeness in the writing is a lead. Agreement between people is measured in the calibration batch.

## Sixty messages on the belt. The rare class is pulled out first.

sixty messages · seven classes · the rare class

The first sixty enquiries of an example batch, embedded in this tab one by one. Before labelling starts, the items nearest each thin class are pulled forward, so the class that matters most has enough labelled examples to measure agreement on.

A class at a few percent leaves a random sample with a handful of items, too few to score. Pulling the nearest candidates first gives it enough.

The small model stands in for yours. In a programme the same pull runs on your model's embeddings.

## Five kinds of label, each with the agreement statistic that fits it.

image class · condition grade · intent · policy · preference

Each first pass runs in your browser from our own storage.

Percent agreement looks high on any task, because agreement by chance is inside it. Each label type is reported with a chance-corrected statistic and the published band to read it against.

Landis and Koch bands for kappa · Krippendorff's thresholds for alpha

- One vehicle, one box, two classes on it.
- 0.61 to 0.80 substantial · Landis and Koch
- One box, graded 0 to 3.
- grades, 0 to 3 · three evidence pins
- 2 against 3 counts as a near miss, 0 against 3 as a full miss
- Parcel says delivered but nothing is at the door.
- delivery status · next damage claim 0.34
- reported per class, the rare class first
- The driver was rude and the box was soaking wet.
- courier conduct · and item damaged 0.45
- 0.80 reliable · 0.67 tentative · Krippendorff
- Two replies to inq_002. Which one goes out?
- rubric clauses · one of them decides
- published labeler agreement on preference data sits near three in four

Rubric · answers the ask, commits to nothing the policy forbids, one next step.

Sorry about that. I have stopped the plan and the last charge is being returned to your card; you will see it within five business days.

Thanks for flagging this. I can see the charge and I have opened a billing review; someone will be in touch.

## Volume starts after three people have read the definitions the same way.

Seven stages and one gate. In the calibration batch three people label the same items independently, their agreement is measured per class, and production volume is committed once the rare class clears it.

The gate · agreement on the rare class, per version, before production

### Taxonomy

- The classes you have, or the ones still being argued about.

### Definitions v1

- One sentence per class, written before labelling starts.

### Calibration batch

- Three people label the same items independently.

### Double labelling

- Two people on every production item.

### Adjudication

- Every split goes to a reviewer, with the rule applied and the reviewer's id.

### Gold

- Items with a known answer, seeded through the batch and scored per annotator.

### Delivery

- The labels and their record, versioned.

## The record is the product.

- **Definitions**: v2 · one rule added
- **Rule**: line 4 · what we must act on first
- **Decision**: refund request · a person
- **Readings**: three clauses · three classes
- **Residency**: EEA · DPA on the engagement

The record for the example message under definitions v2, written on the message it belongs to. Votes and agreement per class are added when the calibration round has run.

EU AI Act Article 10 asks for representativeness, error examination, bias examination and provenance. The definitions in force, the agreement per class and the split log arrive as files, so your assessor reads the history with the labels.

- **Label set**: Every class with the definition in force. The definition file, versioned, with the tie-break.
- **Definition changes**: What changed, when, and what it fixed. The diff between versions and the round that motivated it.
- **Agreement**: Per class, with gold seeded through the batch. Pairwise and pooled kappa, alpha, the rare class first.
- **Adjudication log**: Every split, decided in writing. The decision, the rule, the reviewer's id, and every vote kept.
- **Residency**: EEA under a signed DPA. The processing record and the sub-processor list.

## Bring the taxonomy you have, or the one you are still arguing about.

class set · volume · languages · acceptance

We scope the class set, the volume, the languages and the acceptance plan. When the definitions are still open, writing them is the first work package, and the calibration batch tests them before volume.

An engineer or delivery lead replies inside one EU business day with a feasibility read.

- [Text annotation and preference data](https://ypai.ai/annotation/text-annotation-services/)
- [Model and agent evaluation](https://ypai.ai/data-solutions/evaluation/)
- [AI data annotation](https://ypai.ai/ai-data-annotation/)

A label is a decision. A decision needs a definition, a tie-break and a record of who made it.

## What is the difference between data labelling and data annotation?

Both words are used for both jobs across the market, and we deliver both. What differs is how they fail. A mark fails on position: the edge, the occlusion, the dark frame. A class fails on judgement: an ambiguous definition, two categories that overlap, a taxonomy that drifts between people, a rare class nobody sees. Labelling programmes are scoped on this route, and spatial marking on <a href="/ai-data-annotation/">AI data annotation</a>.

## How does YPAI decide a case that sits between two classes?

The guideline carries a tie-break rule that says which class wins when more than one fits. When two classes still fit after the rule, the item goes to a named reviewer, and the decision, the rule applied and the reviewer's id are written into the adjudication log that ships with the batch.

## How many people label each item?

Two, always, with a third who settles disagreement. Agreement between the two is measured and reported per class. In the calibration batch three people label the same items, so every split has a majority and a minority to read.

## Which languages does YPAI cover?

YPAI supports work across 150+ languages and dialects, with native-speaker reviewers concentrated in European, Nordic and major Asian markets. Lower-resource languages are quoted per project, because reviewer recruitment sets the pace there.

## Do you use models to pre-label?

Where it helps, as model-assisted annotation with human verification. A model's output is a first pass, and a person accepts or corrects every item before delivery. The models on this page do a different job: they read definitions and messages to show where a decision is weak.

## Where is the data processed?

Inside EU/EEA infrastructure under a signed Data Processing Agreement, with lawful basis, purpose limitation and data minimisation set at scoping. Subject-rights requests route through the <a href="/gdpr/request/">data request form</a>.

## How fast does YPAI reply after an enquiry?

We reply inside one EU business day after a submission through the <a href="/contact-us/">contact form</a>, with a feasibility read, the next concrete step and an estimated scope window. The first reply comes from an engineer or delivery lead who has read the brief.
