We start with what we cannot measure
Each rig measures a different set of signals. What buyers actually look at is wrist SE(3) and thumb-to-index aperture, and the minimum setup that produces them is T2. Data captured on T1 cannot be sold.
- T1Practice only
Smart glasses + watch
Setup
- Ray-Ban Meta Gen 2
- Apple Watch ×2
- iPhone (optional)
What it measures
- Egocentric video 3K/30fps
- Wrist acceleration and angular rate 100Hz
- Heart rate
What it cannot measure
- Cannot measure wrist SE(3) — only inertial signals come out
- Finger aperture is estimated from video, so it is weak in low light and under occlusion
- Developer stream limited to 720p · 3-minute clips
Below sellable quality. For securing consent and rights, and for rehearsing the procedure.
- T2Recommended
Headset + IMU suit + gloves
Setup
- Meta Quest 3S
- HaritoraX 2 Pro (8 IMU)
- Rokoko Smartgloves II
What it measures
- 26 hand joints at 60Hz — wrist SE(3) against headset SLAM
- Finger IMU at 250Hz · unaffected by occlusion
- Full-body pose
What it cannot measure
- IMU drift — recalibration every 20–60 minutes
- Hands outside the headset's field of view are lost
- Upper and lower body coordinate frames need aligning
The recommended minimum sellable quality. The baseline setup for the first pilot.
- T3Pro
Professional mocap + EMF gloves
Setup
- Movella Xsens Link (17 IMU)
- Manus Metagloves Pro
- Apple Vision Pro
- Vive Ultimate Tracker ×2
What it measures
- Full body at 240Hz · orientation ±1°
- 25-DoF finger absolute position (EMF, no drift)
- Egocentric video + gaze
What it cannot measure
- No contact force or touch — needs separate hardware
- EMF degrades in outdoor magnetic environments
For research and evaluation sets. Deployed on small, high-quality segments.
What we read from the body
The ninth item is rights because that too is a signal we measure and record. However precise the data, it cannot be sold without a consent scope attached.
| Code | Signal | Unit | Min. tier |
|---|---|---|---|
| WR-01 | Wrist pose | SE(3) 4×4 | T2+ |
| AP-02 | Finger aperture | m (thumb–index) | T2+ |
| JT-03 | Full-body joints | 21 + 40 hands/feet | T2+ |
| EG-04 | Egocentric video | RGB 30fps | T1+ |
| IM-05 | Inertial | 9-axis 100–240Hz | T1+ |
| CT-06 | Contact | grasp start/end | T3 |
| GZ-07 | Gaze | fixation point | T3 |
| PS-08 | Pose source | uint8 0–3 | All |
| RT-09 | Rights | consent scope · withdrawal terms | All |
Glasses alone are not sellable
Hand positions estimated from a monocular camera carry a depth error of 5–11cm. It looks plausible to the human eye, but a robot reaching along those values overshoots the object or collides with it.
That gap shows up as success rate in real robot trials. Training on stereo-tracked hand data gave 95%, monocular estimates gave 45%. Same task, same robot — only the data quality differed.
So we do not use T1 for sellable capture. We use it to refine the consent procedure, validate the task sheets, and let contributors get used to wearing the rig. With that groundwork done, T2 capture passes from the first session.
Source · HumanEgo real-robot evaluation (Aria stereo 95% / WiLoR 45% / HaMeR 32.5%) · rig specifications per manufacturers' published figures, 2026-09