Robotics · Perception and task data

Every frame knows where it stood.Every space said yes.

Stereo, LiDAR and pose collection across gardens, fields, warehouses and homes, task demonstrations from consented workplaces, and annotation with measured agreement, with consent on every record and the calibration your perception team maps to its own rig.

The three questions Where · how precisely · lawfully

A robot that sees is easy. Three things make its data usable

Whether it was collected where the robot will run. Whether the frame knows where it stood. Whether the footage can lawfully exist.

  1. Was it collected where the robot will run?

    DE · aisle · artificial · clutter · people

    A model trained on one climate's lawns meets moss and gravel in another. One trained in one warehouse meets a different floor, racking and light in the next. The programme is specified as a grid of country, ground, light, condition and presence of people, and coverage is counted per cell against that grid, with the denominator showing.

  2. Does the frame know where it stood?

    Pose fixed · stereo pair · calibration cal-214

    Outdoors an RTK position, indoors a pose anchored to the space, and the rig's calibration in every delivery, so the depth in the image and the position on the ground can be checked against each other and mapped to your camera baseline. A float or lost pose stays in the record and never counts as ground truth.

  3. Can the footage lawfully exist?

    Consent c-4471 · space owner · Article 6

    A garden camera sees neighbours, a warehouse camera sees staff, a home camera sees a family. Every record carries the consent of the space's owner and the people in it, third parties are de-identified before an annotator sees the frame, and the material stays in the EEA.

What YPAI collects and labels Four engagements · one robot at a time

Start from the ground your robot has to cross

Collection, task demonstrations, annotation and held-out ground truth, each bought on its own and configured for one robot, one rig and one coverage grid. Four things are settled before the first walk.

Agreed before the first walk

  • The coverage grid, cell by cell
  • The rig, its mount and its calibration
  • Consent, de-identification and residency
  • What the record per frame contains
  1. Collection where the robot will run

    Stereo, LiDAR and pose sequences from the gardens, fields, aisles and homes your robot will actually work in.

    Identity-verified contributors walk a stereo camera rig, with RTK GPS outdoors and an anchored pose indoors, through consented gardens, orchard rows, pavements, warehouses and homes in several European countries. The light, conditions and presence the grid names, hazards staged where the robot will meet them, and back to the same place when the grid asks for it again.

  2. Task demonstrations

    First-person video of the tasks an arm or a humanoid has to learn, recorded in consented workplaces.

    Egocentric video of hand-object interaction, tool use, grasp and place, assembly and task sequences, recorded by identity-verified contributors in consented workplaces with synchronised device and sensor data, and annotated with action labels, temporal segments and keypoints under an action taxonomy written for the task.

  3. Perception annotation

    Masks, keypoints, 3D boxes and fused sensors two annotators agree on, measured on the delivered frames.

    Semantic and instance segmentation, keypoints for pick-and-place and for people and pets, 3D cuboids and per-point segmentation on LiDAR, and sensor-fusion alignment across camera, LiDAR and IMU, drawn under a written task protocol. Two annotators per frame, agreement measured as IoU per class, and a frame below the threshold goes to adjudication instead of being averaged.

  4. Held-out ground truth for evaluation

    Sequences your model has never seen, with the pose and depth to score it against.

    A held-out set built to the same grid and kept out of training: stereo pairs, point clouds, trajectories and calibration, so VSLAM, depth, segmentation and a vision-language-action policy can be scored against a physical record rather than against another model, and a simulation-trained model against ground that exists. The same set reruns on every release.

Scope the first grid

Run the thesis Five axes · one walk · the record per frame

Specify the grid. Walk a cell. Read the record

You are the perception lead. Pick the countries, grounds, light, conditions and presence your robot will meet, walk one cell of that grid, and read the record every frame arrives with. The bench is a toy grid, so the mechanics are visible.

  • The grid is the denominator.

    Coverage is counted cell by cell against the grid you specified, never as a percentage of frames collected.

  • A lost pose is not ground truth.

    Frames with a float or lost pose stay in the record as calibration evidence and are excluded from coverage and from every ground-truth set.

  • Every frame carries its consent.

    The consent record, the contributor, the rig and its calibration travel with the frame, so any sample can be traced back to the space and the day.

Collection bench · Robot programme Grid 48 · Walks 0
  1. Grid
  2. Walked
  3. Recorded
  4. Counted

The coverage grid

Countries
Grounds
Light
Conditions
Presence

Coverage counts a cell only when it holds the frames the grid asks for, from fixed-pose frames, RTK outdoors and an anchored pose indoors. Float and lost poses stay in the record and never count.

Walk a cell

Cell NO/lawn/day/dry/empty

Pick a cell on the grid, then walk it.

    The record per frame

    Frame
    Cell
    Pose
    Hazard
    Consent
    Contributor
    Rig
    Calibration

    Coverage

    Cells covered
    0 / 48
    Fixed-pose frames
    0
    Hazards captured
    0
      A toy grid and a toy rig, so the mechanics are visible. Nothing on this bench is a result of ours.

      Run the second thesis Two annotators · IoU per class · one threshold

      Two masks on one frame. The agreement is measured

      You are annotator one. Three frames from this page's own plates are on the bench, a lawn edge, a warehouse aisle and a kitchen floor, each with floor, boundary and obstacle to draw. Annotator two has drawn them blind. The overlap per class is computed here, one class below the threshold sends the frame to adjudication, and a model can be run on the frame in your browser to see what a first pass looks like next to a person's mask.

      • The second annotator is blind.

        Annotator two never sees annotator one's masks, so the agreement measures two judgements, not one copied.

      • One class fails the frame.

        Floor, boundary and obstacle share edges. A wrong boundary moves the others, so a single class below threshold sends the whole frame to adjudication.

      • The threshold is written down.

        The protocol names the IoU floor for the task, and every verdict is logged against it with both masks kept.

      Task protocol RP-2 Protocol 1.1 · Threshold 0.80
      The frame
      • Annotator one · you
      • Annotator two · blind
      • The model · SlimSAM

      Annotator one · you

      Draw the three masks. Annotator two is revealed when you have drawn.

      Verdict

      1. Floor none
      2. Boundary none
      3. Obstacle none

      Mean IoU none

      The model · first pass

      SlimSAM runs in your browser on this frame, at the marked point. Nothing leaves the tab.

      A model's first pass is marker A. A person is marker B. Neither is a delivered result.

      This session

      Accepted
      0
      Adjudicated
      0
      Mean IoU
      none
      Three plates and toy masks, so the mechanics are visible. Not a labelling tool. The model runs in your browser and nothing here is a result of ours.

      Who we collect for Six kinds of robot · outdoors, indoors, on the line

      Six kinds of robot, each with its own ground

      A mower, a warehouse robot and an assembly arm do not fail in the same places. The grid, the rig and the hazards are written for the ground your robot has to cross.

      1. Robot mowers and yard robots The edge, the hose and the dusk Grass against gravel and beds, staged hoses and toys on the lawn, leaves and frost on the same coordinates, and dusk captures for the hours a mower runs unwatched.
      2. Agricultural field robots Rows, canopy and a crop taxonomy Orchard and vineyard rows, canopy occlusion and mud, with a crop-segmentation taxonomy and a seasonal protocol written per programme.
      3. Sidewalk and delivery robots Kerbs, crossings and passers-by Pavements, dropped kerbs and driveways where the camera meets people and plates, collected under consent and de-identified before annotation.
      4. Warehouse and logistics robots Aisles, pallets and people at work Racking, pallets left in the aisle, spills, dock doors and staff crossing the path, under artificial light and dust, with an anchored pose on every frame.
      5. Home and service robots Floors, thresholds and a family's things Cables, toys, rugs and thresholds in consented homes, pets and children staged on purpose, low light and clutter, and third parties de-identified before annotation.
      6. Arms and industrial vision lines Grasps a line can act on Pick-and-place keypoints, defect masks, task demonstrations for assembly, and navigation segmentation for the robots that run indoors, under the same two-annotator agreement.

      Collection and annotation are configured per engagement. Agricultural programmes add a crop taxonomy and a seasonal protocol; task-demonstration programmes add an action taxonomy written for the task.

      What leaves the bench One row per frame · the fields a reviewer checks

      The record your perception team and your review will read

      Every frame with its cell, its pose fix, the hazard in view, the consent it belongs to, the contributor, the rig and the calibration. Written by the walk you ran above.

      Read a row left to right: the walk and frame id, the cell it was collected in, then the pose fix, the hazard in view and the consent, contributor, rig and calibration it travels with. Counted says whether the frame counts toward coverage or stays in the record only.

      1. No frames yet. Walk a cell on the bench above.
      Frames
      0
      Fixed
      0
      Hazards
      0

      Synthetic record · generated from this session

      The first grid One country · one ground · one rig

      One country, one ground, one rig configuration

      A pilot is the programme above at its smallest honest size. The grid, the rig and the record's fields are fixed before the first walk, one country and one ground are collected and labelled, and the pilot ends with a delivery you can map to your own hardware and a decision. You judge it against your model, not against a sample reel.

      Fixed before the first walk: the coverage grid, cell by cell, the rig, its mount and its calibration, consent, de-identification and residency, the IoU floor below which a frame is adjudicated, and what the record per frame must contain.

      1. Grid

        Countries, grounds, light, conditions and presence, written as cells with a frame count each.

      2. Rig

        Stereo baseline, LiDAR where the robot has it, mount and calibration matched to your rig, with the file in the delivery.

      3. Walks

        Consented gardens, aisles and homes walked cell by cell, hazards staged, third parties de-identified before annotation.

      4. Labels

        Floor, boundary and obstacle masks with IoU per class, adjudication kept in the record.

      5. Decision

        Continue to the full grid, change the grid, or stop, with the record as the evidence.

      The pilot ends with a delivery you can map to your rig, not a reel you have to believe.

      Scope a pilot

      Image 09 · the aisle, the floor line to the pallet

      Consent, third parties and residency GDPR Article 6 · de-identification · EEA

      The controls a review of robot footage will ask for

      A Norwegian company under GDPR. Every record carries the consent of the space's owner and the people in it, third parties are de-identified before annotation, and the material is stored in Europe by default and processed in the EEA where required.

      Jurisdiction
      Norway · GDPR-native
      Lawful basis
      Consent per record · Article 6
      Third parties
      Faces and plates de-identified before annotation
      Contributors
      Identity-verified · EEA-resident
      Storage
      European by default · EEA where required
      Calibration
      Rig intrinsics and extrinsics in the delivery
      Contracts
      Standard DPA terms · SCCs available
      Erasure
      30-day end-of-contract SLA

      Scope a programme A feasibility read · a bounded first grid

      Tell us the robot and the ground

      We scope against the ground your robot has to cross, agree the grid, the rig and the consent model, and prove the method on one country and one ground.

      1. Scope
      2. Grid
      3. Rig
      4. First walks
      5. Delivery

      A bounded first grid carries its own acceptance criteria and its own record. You judge it against your model, not against a sample reel.

      Briefs are treated as confidential. We are used to programmes that cannot be named and footage that cannot leave the EEA.

      Image 11 · the part on the belt, the gripper open above it

      The brief Tell us the robot, where it runs, and your rig and sensors. We reply with a feasibility read.

      What the programme needs (optional)