AI Data & Evaluation · For AI companies and model developers

We are the humans in your loop. 210,000 of them, run like an operation.

Speech and multimodal collection, annotation, demonstration and preference data, evaluation and red teaming. Made and checked by people, in 150+ languages, with lineage on every unit.

made judged, accepted failed, filed in review Manifest 0412 · forming · lineage on every unit

What we deliver

Buy the data. Buy the judgement. Or both, from one floor.

Ten things a lab buys from YPAI. Five make data, five judge it. The record that trains the next checkpoint is the record that proves the last one: same people, same consent, same lineage.

  1. D01 · Data made

    Speech and audio collection

    Read and spontaneous speech in 150+ languages: studio, field, remote and device-based, transcripts verified by native reviewers.

    Speech data Lands as Speech corpus

  2. D02 · Data made

    Image, video and sensor collection

    Collected to your specification, the rare case covered on purpose, consent recorded per contributor.

    Data collection Lands as Annotated image set

  3. D03 · Data made

    Annotation and validation

    Polygons, boxes, attributes, segmentation, transcription; second annotator on sampled items, agreement tracked.

    Annotation Lands as Annotated image set

  4. D04 · Data made

    Demonstration and SFT data

    Instruction and demonstration data written by people qualified in the domain: engineers, mathematicians, lawyers, clinicians.

    AI Data & Evaluation Lands as Demonstration set

  5. D05 · Data made

    Dataset licensing

    Sourcing and rights-cleared licensing where building from scratch is the wrong spend.

    Datasets Lands as Licensed corpus

  6. D06 · Data judged

    Preference data

    Pairs judged against the last accepted checkpoint by calibrated graders, with a written rationale on every choice.

    Evaluation Lands as Preference and red-team set

  7. D07 · Data judged

    Model and multimodal evaluation

    Human grading of what automated metrics read worst: speech, translation, generated image and video, artefact and authenticity review.

    Evaluation Lands as Evaluation record

  8. D08 · Data judged

    Red teaming

    Taxonomies, probe libraries and domain-specific attacks in the languages your users write; every break filed as a regression case.

    Evaluation Lands as Preference and red-team set

  9. D09 · Data judged

    Agent trajectories

    Multi-step task records with tool-use traces, judged at the step where the run failed, not only at the final answer.

    Evaluation Lands as Trajectory record

  10. D10 · Data judged

    Production feedback loops

    Human review of live outputs after ship, returned as graded training and evaluation data for the next run.

    Evaluation Lands as Feedback record

The fab

Your data is built in layers, by people.

Contributors come on in waves, language by language, until the floor holds the coverage your spec asks for. Then the layers go down: capture first, annotation, demonstrations, preference on top.

Nothing is generated into existence; every layer is a person's work, calibrated before production.

  • 210,000+identity-verified contributors
  • 150+languages, native reviewers
  • 50+countries

Network reach, not a headcount on your project: each engagement mobilises the slice it needs, by language, country, dialect, demographic, domain and professional credential.

L4 · Preference Pairs judged against the last accepted checkpoint, rationale written
L3 · Demonstration Instruction and response written by people qualified in the domain
L2 · Annotation Labeled to your schema, second annotator on sampled items
L1 · Capture Recorded by identity-verified people, consent per unit, revocable

YPAI delivers for

  • Cerence AI
  • Nexdata
  • Hyundai
  • BYD
  • Honda
  • Kia

One unit · 0412-0189

Cut one unit open and the people are still in it.

A person made it under a consent they can revoke. A qualified contributor labeled it. A calibrated reviewer checked it and a native reviewer sampled it. When it is a pair, a qualified human judged it and wrote down why.

  1. L4 · Preference Pairs judged against the last accepted checkpoint, rationale written
  2. L3 · Demonstration Instruction and response written by people qualified in the domain
  3. L2 · Annotation Labeled to your schema, second annotator on sampled items
  4. L1 · Capture Recorded by identity-verified people, consent per unit, revocable

L4 · one preference judgement

A · last checkpoint

Yes. A leap year is added every four years so the calendar keeps pace with the seasons. The rule has no exceptions.

B · candidate

Almost. Years divisible by 100 are skipped unless also divisible by 400. So 1900 was not a leap year and 2000 was.

Chosen: B “A states the rule without its exception. B states the exception and gives the test case.”

The test

Then we prove it.

Every engagement ends the same way: the candidate's map laid over the last accepted one. Units your data improved, units that regressed and were filed as regression cases the next checkpoint has to hold. Unit by unit, judged by calibrated humans.

Specimen 0412 · 341 units

Improved
37
Regressed, filed
9
Unchanged
239
Not judged
56

Every regression is a case the next checkpoint has to hold. Preference data, model and multimodal evaluation, red teaming, agent trajectories, production feedback. Evaluation

Candidate against last accepted checkpoint · judged pair by pair

Cut from the same wafer.

What lands in your pipeline

You are not buying hours. You are buying these manifests.

Four kinds of delivery, diced from the same wafer. Made on the left, judged on the right, the same floor on both.

  1. Speech corpus

    Read and spontaneous clips, verified transcripts, consent per contributor

    Made
    Studio, field, remote and device-based collection; transcription verified by native reviewers.
    Judged
    Sampled by an independent native reviewer; accepted against agreed floors; rights cleared for training.

    Lands from D01 · Speech and audio collection

  2. Annotated image set

    Polygons, boxes, attributes to your schema

    Made
    Multimodal annotation and dataset validation, the rare case covered on purpose.
    Judged
    Second annotator on sampled items; agreement tracked; artefact and authenticity review for generated media.

    Lands from D02 · Image, video and sensor collection D03 · Annotation and validation

  3. Demonstration set

    Instruction and accepted response, domain-matched

    Made
    Written by qualified contributors: code by engineers, proofs by mathematicians, contracts by lawyers.
    Judged
    Checked by a second qualified contributor; native review in every language shipped.

    Lands from D04 · Demonstration and SFT data

  4. Preference and red-team set

    Pairs with rationale, findings filed as regression cases

    Made
    Pairs presented against the last accepted checkpoint; probes from taxonomies and probe libraries.
    Judged
    Chosen with a written rationale by calibrated graders; every break rechecked on the next checkpoint.

    Lands from D06 · Preference data D08 · Red teaming

Run like an operation

Data you can audit, delivered in batches you can accept.

Five gates before a batch is yours. The fourth is a person, in every language shipped. Rework until it passes; then it is yours.

Batch 01 Qualification Tested into the task: language, domain, device, credential. 02 Calibration Graded against shared rubrics before production; disagreement resolved, not averaged away. 03 Production Collection, annotation, pairs and probes at volume; agreement tracked throughout. 04 Native review Human A second, independent native reviewer over sampled work in every language shipped. 05 Acceptance Delivery manifests against agreed criteria. Rework until it passes; then it is yours. Verdict
B-0141 Speech corpus, two locales Delivered
B-0142 Annotated images, polygons and attributes Delivered
B-0143 Preference pairs, side-by-side Delivered
B-0144 Red-team probes, regression recheck Delivered
B-0145 Spontaneous speech, two locales In production

4 delivered · 1 in production Acceptance against agreed criteria · manifests versioned · rework until it passes

The edge

The law is the bottom layer, not a sticker.

Every contributor is identity-verified. None of them is shown.

Contributors
Identity-verified, under consent and legal-basis handling
Consent
Recorded per unit, revocable
Lineage
On every unit, from contributor to manifest
Jurisdiction
Norway, under GDPR
Processing
EEA where the engagement requires it; residency and subprocessors per project
Controls
Isolation, client-specific environments, audit records, retention and deletion

Detail: the ethical framework.

Notes on the label · what model teams ask first

Can you deliver in our schema and format?

Yes. Scope defines the schema, the acceptance criteria and the delivery manifest before production. Batches are versioned and reworked until they pass; then they are yours.

How are graders calibrated?

Contributors are tested into the task by language, domain and credential, then graded against shared rubrics before production. Agreement is tracked throughout and an independent native reviewer samples every language shipped.

Where is the data processed?

YPAI operates from Oslo under GDPR. EEA processing where the engagement requires it; residency, subprocessors and transfer controls are defined per project.

Do you red-team, and what do we get back?

Yes. Taxonomies, probe libraries and domain-specific attacks in the languages your users write. Every break is filed as a regression case with a recheck against the next checkpoint.

One decision

The next checkpoint ships with proof, or it ships with hope.

Bring the modality, the languages, the volume, and what you need to prove. YPAI scopes the slice of the floor, the rubric, and the first batch.

  1. Scoping call
  2. Pilot batch
  3. Production
Your wafer · each field you complete starts a ring
Modality (optional)