The aeffer-v1 schema
What we store is joint SE(3) and sensor origins. The point cloud you see in the viewer is only a visualisation unfolded from those parameters — we do not store it. Point count is not a quality measure.
The structure of a single frame
| Path | Type | Description |
|---|---|---|
| schema_version | string | Always "aeffer-v1" |
| task_description | string | Task name (matches the task sheet verbatim) |
| actors[] | array | Role (demo/subject) · body profile · height |
| capture_rig | string | Rig identifier (t1 · t2 · t3 · …) |
| fps | uint16 | Sampling rate |
| frame_index | uint32 | Frame number |
| timestamp_ns | uint64 | Nanoseconds from session start |
| pose_source | uint8 | 0 measured · 1 monocular estimate · 2 interpolated · 3 synthetic |
| transforms/<joint> | float32[4][4] | Joint SE(3) homogeneous transform |
| confidences/<joint> | float32 | Per-joint confidence 0–1 |
| sensors/ | object | Sensor origin coordinates · count (surface points are not stored) |
| derived/aperture_m | float32[2] | Left/right thumb-to-index distance (m) |
| derived/contact | uint8[2] | Left/right grasp state |
| rights/consent | object | Consent scope · withdrawal terms (required) |
71 joints
| Region | Count | Composition |
|---|---|---|
| Torso | 5 | hip · spine · chest · neck · head |
| Arms | 8 | shoulder · elbow · wrist · hand (each side) |
| Legs | 8 | hip · knee · ankle · foot (each side) |
| Fingers | 40 | 5 fingers per hand × MCP · PIP · DIP · tip |
| Toes | 10 | 5 tip points per foot |
Two things this schema guarantees
- PS-08
We do not mix measured and estimated values
Every frame carries
pose_source. Buyers can filter out estimated segments before training, and we can put a number on quality. Without this field there is no way to verify how far the data can be trusted. - /derived/
We isolate derived values
Computed values such as aperture and contact sit separately under
/derived/. They never mix with the original measurements, so a change to the formula cannot contaminate the source — and values derived from differently licensed models (MANO and the like) stay separable.
It goes out as LeRobot v3
The de facto standard in robot learning is the LeRobot v3 format (parquet + mp4). aeffer-v1 is designed to convert into it without loss, so buyers do not have to change the pipeline they already use.
The format is not a moat. We are not inventing a proprietary format to lock anyone in — going out as a standard is what makes it sellable. The value is in the quality of the data and the rights attached to it, not the container.
See a real frame dump in the viewer →Quality criteria →Licence →