The data and evaluation work behind AI systems that hold up in production.
YPAI runs managed collection, annotation, validation and human evaluation across speech, video, image, text, agents and robotics. Every delivery carries its record.
Identity-verified contributors · Credentialed domain experts · 150+ languages · 50+ countries · EEA residency by default
Move across the plate to scrub the run
Four problems. Four kinds of work.
The data you need does not exist.
- modalities
- speech, video, image, text, sensor, interaction
- settings
- studio, field, in-car, remote, moderated, on site
- reach
- 150+ languages, 50+ countries
- sourcing
- existing data licensed first, only the gap collected
- Specialised collection
- Speech dataLow-resource languagesConversational and multi-speaker audioIn-car and far-field audioVoice recordingVideo dataOn-camera speech videoEgocentric and robot demonstration videoPhysical AI dataSensor and LiDARImage collectionHuman motion and gestureExpert-written text and promptsHealthcare collectionGeospatialSynthetic and augmented dataReady-made datasets
Your data exists but lacks structure.
- labels
- boxes, segments, transcripts, events, keypoints, rankings
- controls
- guideline, gold tasks, agreement, adjudication, sampling
- review
- second labeller and adjudication before a label counts
- traceability
- every decision bound to the guideline it was made against
- Specialised annotation
- VideoImageTextAudio and speechLiDAR and 3DSensor fusionSegmentationKeypoints and poseObject tracking and temporal eventsMedical imagingNamed entitiesTranscription and diarisationDocument extractionRobotics visionPreference and RLHF dataAgent trajectoriesContent moderationOntology and guideline design
You cannot trust the data you have.
- tests
- conformity, representativeness, duplication, contamination, rights
- sampling
- statistical lots and gold sets
- decision
- accept, remediate or replace
- record
- every failed item named, with its reason
- Specialised validation
- Audio data QASpeech specificationsProvenance auditDataset auditsLabel QA and agreement analysisRepresentativeness and bias reviewDuplication and contamination checksConsent and rights verificationAcceptance sampling and gold setsSensor calibration and alignmentSynthetic-to-real validationVideo integrity and synthetic mediaArticle 10 data governanceEEA data residency
sampled lot · 320 items 3 items → remediation
You cannot trust what the model does.
- grades
- response grading, preference, red-teaming, regression
- against
- your rubric, population, languages, thresholds
- reviewers
- credentialed domain experts where the work requires them
- before
- a release decision
- Specialised evaluation
- Clinical model evaluationSpeech evaluation projectASR benchmarkResponse grading and rubric designPreference and pairwise comparisonRed teaming and safetyAgent and tool-use evaluationRetrieval and RAG evaluationJudge calibration against human gradingMultilingual model evaluationFactuality reviewRegression testingCoding evaluationSTEM and mathematicsLegal and finance expertsComputer-vision evaluation
response B meets the threshold · response A fails on source use
Selected clients
Every decision stays bound to the frame.
- captured consent recorded · 48 kHz · 24-bit · far-field Proprietary video and audio collection platforms with verified per-contributor consent
- labelled 3 objects · 1 region · adjudicated against the guideline Native human-review console as the platform core
- inspected sampled lot · coverage gap · 3 items remediated Self-hosted European annotation servers
- evaluated response B accepted on the rubric Credentialed domain reviewers where the work requires them
- delivered manifest · integrity check · erasure within 30 days Residency, subprocessors and transfers defined per engagement
Every delivery carries the records the engagement requires and an integrity check on what is delivered.
One pilot against your requirement.
- scope
- your specification and acceptance criteria
- terms
- scope and commercial terms agreed before it starts
- review
- against the agreed criteria, in a pilot workspace
- production
- a separate decision, taken after the review
- Agreed before the first night
- Scope of workTechnical requirementsAcceptance criteriaData protection and rightsCommercial structureRemediation and change control
Production: a separate decision, taken after the pilot review
The model is one component. YPAI builds the working system around it.
- system
- assistants, agents, document workflows, integrations
- failure
- comes back here as the next data requirement
- team
- the same team builds the system and the data
- release
- the record travels with the decision
- From a production failure
- Discuss a targeted project
Use model failures to define the next data workstream.
Scope a data or evaluation brief.
Bring the use case and what you already have: modality, intended use, volume, languages or markets, format, devices or environments, deadline, existing data, acceptance criteria, and any processing, rights or security requirements. YPAI comes back with what has to be clarified before scope, price and terms can be agreed.
- EU AI Act
- For high-risk AI systems under the EU AI Act, the same delivery records map to what Article 10 expects buyers to hold: dataset origin, collection method, representativeness and documented bias review. How this maps to AI Act risk classes
- Reply
- Reply inside one EU business day with a feasibility read.
Your Personal AI AS · Lysaker, Norway