TRACKING / VOS / ACTION / POSE
Video annotation,tracked across every frame.
Multi-object tracking, video object segmentation, action localization, pose, re-ID. SAM 2 assisted, kappa-gated, HOTA-reported per delivery.
HOTA + IDF1 reporting EEA-resident
A clean frame can hide a broken track.
A bounding box can be correct at one timestamp while the sequence fails around it.
The same object may disappear behind an occluder, leave the field of view, return under different lighting or move into another camera. A useful annotation must preserve the right identity, boundary and event state throughout.
Review the sequence as a track, not as a stack of images.
What the readout exposes
Identity continuityDoes the same object keep the same track ID through visibility changes?
Track record
Window f 0182 - f 0271 · 90 frames · single cameraThe failure often appears between two individually correct frames.
For regulated video, the annotation record is part of the delivery.
Ontology, de-identification, reviewer access and change history are not administrative details added after annotation. They determine whether the resulting data can be understood, reviewed and used within the intended workflow.
The annotation specification can document:
- class definitions
- source and provenance fields
- sampling decisions
- annotation and review methods
- known limitations
- version and change history
These records can support the customer’s own data-governance and technical documentation.
Where footage contains faces, licence plates, patients, employees or other identifiable people, the engagement defines:
- what must be retained
- what must be removed or transformed
- who may access the material
- which reviewers require restricted access
- how exceptions are recorded
- what may be delivered downstream
The method is selected against the intended use and residual risk. The project does not assume that one generic de-identification method is sufficient for every video type.
Where YPAI processes personal data on the customer’s behalf, the engagement can define:
- controller and processor roles
- approved environments
- reviewer permissions
- subprocessors
- transfers
- retention
- deletion
- incident and audit responsibilities
EEA-based processing is available where required. Residency, transfers and subprocessors are established per project.
Five annotation primitives. Five different definitions of correct.
The primitive determines what is labelled, how temporal consistency is reviewed, which metrics are meaningful and what the final delivery contains.
Same footage, different truth condition
Multi-object tracking
The annotation follows the object through visibility changes and retains its persistent identity.
-
What is annotated Objects are located frame by frame and assigned persistent identities.
The hard part Identity must survive partial visibility, full occlusion, scene exit and re-entry.
Typical output Bounding boxes, track IDs, visibility states and frame-level attributes.
Relevant review Track continuity, ID switches, fragmentation, HOTA and IDF1 where applicable.
Common uses Vehicles, people, equipment, animals and other moving objects.
-
What is annotated Each object receives a pixel-level mask across the sequence, linked to a continuous identity.
The hard part The mask must preserve both geometric accuracy and instance continuity as shape, lighting and visibility change.
Typical output Per-frame masks, object IDs, visibility states and exception records.
Relevant review Mask overlap , boundary quality, missed objects and identity continuity.
Common uses Fine-grained object boundaries, surgical scenes, content editing and foreground extraction.
-
What is annotated Actions or events receive a class and a temporal start and end point.
The hard part The first visible sign of an action, full action onset and task completion are different boundaries.
Typical output Temporal segments, action classes, actor links and overlapping-event rules.
Relevant review Boundary tolerance, class accuracy and mean average precision at the agreed temporal overlap thresholds.
Common uses Surgical phases, sports actions, industrial tasks, safety events and content understanding.
-
What is annotated Joint keypoints are linked across frames and associated with persistent person identities.
The hard part Keypoints must remain anatomically coherent through motion, occlusion, truncation and viewpoint change.
Typical output Keypoints, skeletons, visibility flags and person track IDs.
Relevant review Keypoint accuracy , missing joints, identity continuity and temporal consistency.
Common uses Sports, rehabilitation, animation, ergonomics and human activity analysis.
-
What is annotated The same person or object is linked across non-overlapping cameras or separated sequences.
The hard part Clothing, angle, lighting, camera quality and elapsed time may all change while identity must remain consistent.
Typical output Cross-camera identity links, handoff events, timestamps and exception records.
Relevant review Identity consistency , false links, missed links and project-specific retrieval or tracking metrics.
Common uses Multi-camera operations, retail dwell analysis, sports and authorised security workflows.
A project can combine primitives when the model requires more than one form of temporal truth.
Scope the annotation primitiveThe ontology changes with the footage.
Automotive, surgical, surveillance, sports and retail video may all use tracks, masks or event labels. They do not use the same object definitions, privacy boundary or review logic.
Recognise the footage first. Define the annotation contract from its operating reality.
-
Automotive
Annotated Drivers, occupants, gaze, gesture, vehicles, pedestrians, cyclists, lanes and temporal events.
Operating frame In-cabin and outward-facing footage require different identity, visibility and scene rules. Privacy and safety requirements are scoped against the actual project.
Reference structures Driver monitoring, occupant monitoring, multi-object tracking, action and gesture taxonomies.
-
Surveillance and security
Annotated People, objects, movement, cross-camera identity, zones and events.
Operating frame Lawful basis, access, retention, signage and proportionality must be resolved for the intended deployment. A DPIA may be required for systematic monitoring.
Reference structures MOT, re-identification, event windows, entry and exit states, camera handoffs.
-
Surgical video
Annotated Tools, anatomy, surgical phase, actions and interactions.
Operating frame Health data, patient identifiers, reviewer qualifications and restricted access are defined before annotation begins.
Reference structures Phase, tool and anatomy taxonomies. Public research datasets may inform schema design, subject to their access and licensing restrictions.
-
Sports analytics
Annotated Players, ball, pose, actions, formations, possession and spatial events.
Operating frame Footage rights, athlete data, league agreements and intended commercial use must be established for the engagement.
Reference structures Tracking, pose and action taxonomies such as those used in sports-research benchmarks.
-
Retail and content
Annotated People, queues, dwell zones, movement, interactions, shots and content events.
Operating frame The project may require a legitimate-interest assessment, defined retention and a documented moderation or loss-prevention purpose.
Reference structures Dwell events, queue states, zone entry and exit, shot boundaries and content labels.
Production starts at one gate.
Schema, guideline and calibration all serve the same decision:
Is the annotation contract stable enough to scale?
-
Schema and sensitive-content policy
Define:
- class taxonomy
- identity continuity rules
- sensitive-content handling
- required attributes
- output structure
- prohibited or excluded labels
Output Versioned schema and sensitive-content policy.
-
Annotation guideline
Resolve:
- occlusion
- partial visibility
- truncation
- scene exit and re-entry
- overlapping objects
- action boundaries
- ambiguity
- escalation rules
Output Versioned annotation guideline.
-
Calibration round
A shared subset is annotated and reviewed before production.
The round is used to expose:
- unclear instructions
- missing edge cases
- inconsistent object-identity handling
- weak class boundaries
- metric mismatch
- reviewer disagreement
Disagreement feeds back into the schema and guideline.
Output Calibration record and revision decisions.
-
The gate
Where the metrics apply to the task, the reference production gate is:
- Kappa Project-defined
- IDF1 0.75+
- HOTA 0.65+
The statement of work confirms which metrics apply and may define stricter or task-specific thresholds.
Round oneOne or more applicable requirements remain below the gate.
Verdict REFINE
The annotation contract returns to schema, guideline or calibration.
Round twoAll applicable requirements clear the agreed gate.
Verdict PRODUCTION
Only then does production open.
-
Production
Production may combine keyframe annotation, assisted tracking, interpolation and human review where the method fits the footage and acceptance plan.
Assistive tooling accelerates appropriate parts of the work. Correctness is decided at the agreed review gate.
Output Annotated video, tracks, masks, events or keypoints in the agreed format.
-
QA, adjudication and delivery
Production QA can include:
- sample review
- track-continuity checks
- ID-switch review
- boundary review
- class-specific checks
- disputed-item adjudication
- exception handling
- versioned exports
Output Accepted delivery, quality record and known-limitations statement.
One gate. One recorded decision. A visible route back when the contract is not ready.
Define the calibration gateThe annotation is the sequence, not the screenshot.
A production track records what happened between the selected frames.
TRACK CONTINUITY RECORD
track-continuity.json
EVENTS RECORDED FOR REVIEW
-
Entry
The object becomes observable and receives its first project-defined identity.
-
Partial occlusion
Visibility changes, but the identity and boundary rules remain active.
-
Full occlusion
The object is temporarily unobservable. The record distinguishes occlusion from confirmed exit.
-
Re-entry
The object returns. The reviewer confirms whether the earlier identity remains valid.
-
ID-switch event
The system or annotator links the object incorrectly. The exact transition is flagged for correction or adjudication.
-
Action boundary
The sequence records where the defined event begins, continues and ends.
-
Human review
Low-confidence or ambiguous transitions move to a reviewer under the project’s escalation rules.
Where assistive tooling proposes masks or tracks, human review focuses on the cases that determine acceptance: occlusion, identity continuity, boundary drift, scene handoff and ambiguous events.
What the sequence can produce
- persistent track IDs
- frame-level boxes or masks
- visibility and occlusion states
- event intervals
- camera-handoff records
- review and adjudication events
- versioned corrections
Benchmark schemas can enter the project. Their licences do not.
Public benchmarks can help define taxonomy, output format and evaluation.
Research access, annotation licences and familiar file formats do not automatically grant the right to use the underlying footage for commercial model training or deployment.
YPAI treats benchmark structure and source-data rights as separate questions.
-
MOTChallenge
Multi-object tracking Frame-level bounding boxes and persistent track IDs
Tracking format and evaluation alignment
Research and non-commercial terms must be checked against the intended use.
-
DAVIS
Video object segmentation Per-frame instance masks
Mask and temporal-consistency evaluation
Public research access does not establish commercial training rights.
-
YouTube-VOS
Large-scale video object segmentation Instance masks across sequences
Taxonomy and VOS evaluation reference
Underlying video rights and annotation rights must be assessed separately.
-
Kinetics-700
Action recognition Clip-level action classes
Action-taxonomy reference
An annotation licence does not necessarily confer rights to the source video.
-
AVA
Action localisation Per-frame boxes with atomic-action labels
Spatiotemporal action evaluation
Source footage, annotations and downstream use must each be checked.
-
Cholec80
Surgical workflow Surgical phases and tool labels
Clinical-video taxonomy reference
Research, clinical-access and patient-data restrictions remain outside a generic annotation licence.
Nothing crosses the rights boundary because the schema looks familiar.
Every source, licence and permitted use is assessed for the engagement.
The evidence package is defined before production.
The delivery is more than labelled frames.
The statement of work identifies the records the customer will receive, how they relate to the data and which technical or governance questions they answer.
-
Annotation contract
- class schema
- annotation guideline
- object-identity and tracking policy
- temporal-boundary rules
- occlusion and partial-visibility policy
- version and change history
Representative outputguideline.v3.pdf
-
Calibration and quality record
- calibration sample
- applicable metrics
- threshold decisions
- disagreement analysis
- revision actions
- production-gate verdict
hota-idf1.csv
-
Track-continuity record
- track lengths
- ID switches
- fragmentation
- occlusion and re-entry exceptions
- cross-camera handoff results where applicable
track-continuity.json
-
Sensitive-content and de-identification record
- project-specific policy
- excluded or transformed fields
- access restrictions
- exceptions
- reviewer decisions
- residual limitations
-
Processing and delivery record
Contains, where applicable- roles and instructions
- approved environment
- subprocessors and transfer decisions
- delivery manifest
- dataset version history
- retention and deletion terms
- known limitations
article-30-records.pdf
The exact evidence package is defined in the statement of work and applicable data-processing terms.
These records support the customer’s own technical, procurement and regulatory review.
Bring the footage, the model objective and the acceptance problem.
A short brief is enough to start.
YPAI will identify the appropriate annotation primitive, the open ontology decisions, the data and rights boundary and the first calibration step.
- the model or product objective
- sample footage, when available
- the video domain
- the objects, actions or identities of interest
- the expected volume and volume unit
- required output format
- sensitive-content considerations
- existing ontology or benchmark references
- acceptance criteria
- target schedule
- feasibility assessment
- recommended annotation primitive or combination
- open ontology and guideline decisions
- proposed calibration method
- applicable quality metrics
- reviewer and adjudication model
- data-processing and rights requirements
- delivery structure
- evidence-package proposal
- schedule and price basis
- proposed first gate
We reply within one business day with the questions that must be resolved before a pilot or production scope.
Answers that belong in the first conversation.
How do you handle faces, licence plates and patient identifiers?
The method is defined before annotation begins.
Depending on the intended use and residual risk, the project may use exclusion, masking, replacement, restricted access, metadata removal or a reduced representation such as skeleton-only output.
Blur is one method. The method is chosen against the intended use and residual risk.
The chosen method, exceptions and remaining limitations are documented for the engagement.
Can YPAI handle multi-camera re-identification?
Yes, when the engagement defines:
- camera topology
- entry and exit states
- handoff rules
- identity attributes
- time-gap assumptions
- allowed ambiguity
- reviewer escalation
- evaluation criteria
The delivery can include cross-camera identity links, handoff events, false-link review and project-specific continuity reporting.
Do you use SAM 2 or other assisted annotation tools?
Assistive segmentation and tracking can be used where they fit the footage, ontology and acceptance plan.
The tool does not define correctness.
Human reviewers remain responsible for the agreed checks, especially around occlusion, identity continuity, boundary drift, scene changes and ambiguous events.
Tooling is confirmed during scoping rather than promised as a fixed method for every project.
Where is the video processed?
The engagement defines the processing environment, access model, residency, subprocessors and transfer controls.
EEA-based processing is available where required.
Customer-controlled environments or other restricted workflows can be assessed during scoping, but feasibility depends on the technical, security and operational requirements.
How is annotation quality measured?
The metric follows the annotation job.
Tracking may use HOTA, IDF1, ID switches and fragmentation.
Segmentation may use overlap and boundary measures.
Action localisation may use mean average precision at temporal overlap thresholds.
Pose may use keypoint metrics plus identity continuity.
No universal accuracy percentage is meaningful across all of these tasks.
What happens when annotators disagree?
Disagreement is recorded, not silently overwritten.
The item may move to:
- second-pass review
- senior review
- specialist review
- adjudication
- guideline revision
- schema revision
- another calibration round
Recurring disagreement is evidence about the annotation contract, not merely worker performance.
Which output formats can you deliver?
The exact format is defined in the statement of work.
Common structures can include:
- MOT-style tracks
- COCO-style boxes, masks or keypoints
- frame-level and sequence-level JSON
- JSONL or CSV records
- temporal segments
- mask encodings
- CVAT-compatible exports
- customer-defined schemas
Every delivery includes a manifest that identifies the agreed files, versions and fields.
Can YPAI work in our existing annotation environment?
Potentially.
The engagement must define:
- supported import and export formats
- user and reviewer access
- customer security controls
- audit logging
- data movement restrictions
- integration requirements
- responsibility for tooling and support
YPAI can assess a customer-controlled environment or propose a managed workflow.
Start with the sequence your model must understand.
Bring the footage, annotation objective, classes, output requirements and acceptance problem.
YPAI will return the proposed project scope, the open decisions, the calibration plan and the evidence package required to evaluate the first delivery.
One business day to the first scoping response. Production begins only after the annotation contract and applicable acceptance gate are agreed.
FAQ
Frequently asked questions
What video annotation tasks does YPAI deliver?
YPAI delivers object detection and tracking across frames, instance segmentation with temporal IDs, action recognition and event-segment labelling, multi-camera re-identification, 3D tracking with sensor fusion, pose and gesture annotation, and dense per-frame keypoint markup. The deliverable is structured per-frame and per-track so downstream training pipelines can choose the granularity they need without re-running annotation.
Which video formats and resolutions does YPAI handle?
Standard codecs (H.264, H.265, VP9, AV1) and container formats (MP4, MOV, MKV) are supported. Working resolution is matched to the customer pipeline rather than fixed at 1080p, so 4K vehicle footage, drone imagery, and surveillance-grade lower-resolution streams are all handled. Time-stamped sensor traces (LiDAR, radar, IMU) can be annotated synchronously with video for autonomous-vehicle engagements.
How does YPAI handle long-form video annotation?
Long-form video is split into review-friendly segments with frame-level cross-segment validation so track IDs remain consistent across the cut points. Annotators specialise by domain (automotive scene-understanding, retail loss-prevention, industrial vision, surgical recording) rather than rotating through arbitrary content. IAA on tracking metrics (HOTA, MOTA) is reported per batch.
Can YPAI annotate autonomous-vehicle perception data?
Yes. Multi-camera, LiDAR, and radar fusion annotation for ADAS and autonomous-vehicle perception stacks is a core competency. The delivery includes project-specific taxonomies, traceability, edge-case documentation, and EU AI Act Article 10 evidence for the customer audit trail. See automotive solutions for the wider engagement model.
What output formats does YPAI deliver for video?
Deliveries are shipped in COCO-Video, CVAT, KITTI tracking format, custom JSON with frame and track schemas, and any customer-specified output. Frame-accurate timestamps and sensor-trace synchronisation are preserved through the export. The delivery bundle includes a calibration record where multi-sensor alignment was performed.
How are video annotation engagements priced?
Pricing is per-project after scoping. Drivers include hours of source video, target IAA, label density per frame, and any multi-sensor fusion requirements. A flat rate per video-minute rarely reflects real cost. The scoping call at /contact-us/ produces a written proposal with the cost drivers itemised so the customer can see which dimensions trade off against the others.
Is video annotation GDPR-compliant for footage containing people?
Yes. Footage with identifiable individuals is processed under GDPR with documented lawful basis. Where the use case requires it, YPAI runs blurring or face-obfuscation passes on identifiable subjects who did not provide consent for downstream model training. The decision tree is documented in the SOW and the DPA covers the processing terms.