Managed annotation built around the errors your model cannot afford.
Bring one annotation workload. You get back accepted, versioned ground truth, judged against criteria you agreed before any volume started.
A label is not correct because someone submitted it.
- boundary A bounding box identifies the right object and excludes the part the model needs.
- class Two annotators apply the same class to different events.
- speaker Correct words land on the wrong speaker.
- span A span boundary drifts across a shift.
- rubric A rater picks the better of two responses and the rubric behind the choice disappears.
Five failures, five different consequences. One accuracy percentage hides all of them.
You are buying a managed annotation operation for one named workload.
The engagement starts from an accepted unit and annotation contract, and ends in versioned labels and delivery evidence. Between those two points YPAI defines the annotation unit, the error classes, the review method and the acceptance evidence before production volume grows.
One decision holds the rest together. YPAI fixes one accepted unit across pricing, production, quality and acceptance before the first production batch opens.
priced per clip
1
billable unit
accepted per tracked object per frame
24
acceptance units
Get that wrong and the project stops being measurable. Price a video job per clip while errors are counted per tracked object per frame, and you and the delivery lead read the same dataset as two different results.
The same decision locks redundancy, specialist staffing, model assistance, the threshold, cadence, and who remediates. All of it is settled in the contract.
The work runs on YPAI's own platform.
Review, rubric execution, attribution, adjudication and QA all run in YPAI's own human-review console.
YPAI also operates proprietary video and audio collection platforms, with automated ingestion, conversion, QC and structured delivery on the video side, and a controlled participant flow on the audio side. Annotation servers are self-hosted in Europe.
Self-hosted CVAT and Label Studio attach as adapters where a project needs them.
- original Tight box. Trailer excluded.
- finding The extent that carries the meaning is missing.
- correction Box includes the attached load.
- reason Guideline: include connected geometry.
Nothing scales until the guideline survives real material.
This way of running the work needs a boundary, an objective for the model that will consume the labels, and a quality bar you can put in a contract.
Given those, calibration comes before volume. Annotators and reviewers label the same controlled material independently.
- Does the ontology cover the source data?
- Does the guideline produce consistent decisions?
- Can the intended team perform the task?
- Is the quality bar achievable on this material?
A failed calibration gate blocks production scale.
calibration gate, multi-object tracking
Detection turns on recall and boundary tolerance. Transcription is scored as word error rate, or character error rate where the script demands it. Judgement work is scored on rubric compliance and preference consistency. The metric follows the error being controlled. Far-field material carries its own threshold, set during scoping.
YPAI staffs and supervises the annotators, raters, reviewers, adjudicators and specialists for the workload. A person joins a project only after qualifying against the calibrated task. The pool behind that is a contributor network of 210,000+ people across 50+ countries, so a project can be staffed for the language, locale and domain the material needs.
Model-generated pre-labels enter production as candidates. The acceptance and sampling plan decides which of them a person confirms, corrects or annotates from scratch. A confidence threshold alone does not make a pre-label accepted ground truth.
Acceptance is your decision.
You measure each batch against criteria the annotation contract fixes before production starts. YPAI reports against those criteria, and a batch outside the agreed thresholds is remediated or held back rather than released as accepted.
The structural gate first
The contract sets which schema, identifier, relationship and manifest conditions a record must satisfy before it reaches human quality review. Those checks run across the whole batch, not a sample.
Then the arguments that remain
Review finds defects. Adjudication settles the cases where two defensible readings remain. The record keeps the original, the finding, the correction and the reason.
Ontology change is a production event. YPAI versions it instead of editing an active project underneath completed work.
What an accepted delivery carries.
Versioned labels, the ontology, the guidelines, the calibration evidence, the review and adjudication records, the task-specific quality results and a delivery manifest.
So you can ask which specification produced a label, whether it was model-assisted, whether it was reviewed, whether it was disputed, and what changed between versions. Any label traces back to the specification, the review and the version it came from.
The annotation contract records where the work is performed and which roles may access the material. Rights to annotate and to use the result downstream, provenance, source lineage, retention and deletion are defined in the engagement document.
YPAI is an EEA-based supplier operating under GDPR, aligned with Article 10 of the EU AI Act on data governance for high-risk systems, and produces the documentation that supports your own conformity assessment.
Bring the workload.
Tell us the workload, what the source material looks like, the labels you need, and the errors you cannot ship. Send the ontology if you have one, plus how you want quality measured and delivered.
A project lead reviews the workload and replies with the annotation contract it would need.
FAQ
Frequently asked questions
What does a managed annotation operation include?
For one named annotation workload, YPAI defines the ontology, calibration, reviewer operation, quality checks, adjudication, acceptance criteria and versioned delivery before production volume grows. This route explains that operating model. For modality-specific services, see AI data annotation.
What inter-annotator-agreement target does YPAI enforce?
IAA targets are set per task and per class against the brief, not asserted as a fixed company number. Cohen kappa is reported for classification; IoU and mAP for bounding boxes; Dice for segmentation; OKS for keypoints; HOTA and MOTA for tracking. Classes that under-perform the target trigger re-calibration before bulk progress continues. The IAA report ships with every delivery.
How does YPAI document data provenance?
Every record ships with a provenance trail: collection or supply source, date, locale where applicable, reviewer pool, consent identifier where personal data is involved, and any post-processing applied. The bundle is structured for EU AI Act Article 10 dataset documentation so it slots into a customer Technical File. Provenance is part of the deliverable for every engagement.
Can YPAI annotation work plug into our active-learning loop?
Yes. Where the customer has a working baseline model, YPAI runs human review over model pre-predictions and prioritises the lowest-confidence and most-impactful slices. The loop feeds back into retraining batches. Integration patterns (synchronous API, asynchronous batch queue, customer-hosted model) are agreed during scoping.
How are taxonomy changes handled mid-engagement?
Taxonomy changes are scoped as change orders. When the customer adjusts class definitions, YPAI re-runs a calibration batch against the new spec, computes IAA on the changed classes, and either re-labels the affected backlog or stamps it for re-review depending on impact size. Cost and timing impact are quoted before re-labelling starts. How a class is written, and who decides when two classes overlap, is covered on AI data labeling.
How is annotation reviewer expertise documented?
Reviewer qualifications are documented per engagement: clinician credentials for medical work, legal practitioner credentials for contract extraction, native-speaker verification for language work, automotive-specialist experience for in-cabin and AV perception. The qualification record ships with the deliverable so the customer audit trail covers reviewer competence, not only annotation output.