aeffer/ko-household-50h
Capture planned Q4 2026In from the human body,
out to the robot body.
Aeffer turns human demonstrations into a form humanoids can learn from. We gather afferently, and we send efferently.
What sits between the two is us.
See a synthetic sample first
A procedural sample built on the same schema as real captures. Switch tasks, body profiles and close-ups to see which signal comes from where.
pose_source=3 · SYNTHETICThis is a procedurally generated synthetic sample. It is not captured data.
There is no shortage of video. What is missing is rights and hands.
What blocks robot learning is not the volume of data. It is the right to sell it, and the resolution to see hands.
- ( GAP 01 )
The main public datasets cannot be sold
EgoDex and Ego4D (egocentric video) and AMASS (motion capture) all carry non-commercial licences. To train a commercial model, you need separate data with verified wearer consent.
0 / 3Allow commercial training (major public sets checked)EgoDex · Ego4D · AMASS licence terms (checked 2026-09)
- ( GAP 02 )
The hands are not visible
What a buyer actually looks at is wrist SE(3) and the thumb-index aperture. A monocular camera has 5–11cm of depth error, so it cannot produce those values.
95% vs 45%Real-robot success rate · stereo vs monocularHumanEgo real-robot evaluation — Aria stereo 95% · WiLoR 45% · HaMeR 32.5%
Compare the same motion yourself - ( GAP 03 )
Korean homes are hard to find in the data
Floor-seated living, Korean kitchen layouts, kimchi refrigerators, waste separation. We found no demonstration data stated to be captured in homes like these. Even truelabel's Seoul entry is industrial SCARA teleoperation.
0Demonstration data stating a Korean home setting (as checked)truelabel.ai ticker & environment page (checked 2026-09-17/18) · AI Hub hand-arm grasp-manipulation data (2022)
This is what happens when the hands are not visible
pose_source=1 monocular_estpose_source=0 measuredthe robot loses that stretch
A synthetic visualisation. Data leaves the person, aligns at the centre and flows to the robot; when the data for a Korean kitchen task stops, the robot loses the motion and fades. It catches up once the data resumes.
The way in and the way out
We carry from people, and we carry to robots. QA sits between the two.
A contributor puts on a rig
You pick the rig that suits the task, from the three tiers T1, T2 and T3. Consent is taken before the rig goes on. The scope of that consent is recorded: whether commercial training is allowed, third-party transfer, and the terms of withdrawal.
See rig specsWhat we read from the body
Nine signals, each with a code. The last two are what make us different — the provenance of a pose, and rights, are signals we measure too.
| # | Signal | Code | Description | Tier |
|---|---|---|---|---|
| 01 | Wrist pose | WR-01 · AF-SIG | SE(3) 4×4 transform, both wrists | T2+ |
| 02 | Finger aperture | AP-02 · AF-SIG | Thumb-to-index distance (m) | T2+ |
| 03 | Full-body joints | JT-03 · AF-SIG | 21 joints + 30 for hands and feet | T2+ |
| 04 | First-person video | EG-04 · AF-SIG | Head-mounted RGB | T1+ |
| 05 | Inertial | IM-05 · AF-SIG | 9-axis IMU, 100–240Hz | T1+ |
| 06 | Contact | CT-06 · AF-SIG | Grasp start/end, binary | T3 |
| 07 | Gaze | GZ-07 · AF-SIG | Fixation point (Aria / Vision Pro) | T3 |
| 08 | Pose source | PS-08 · AF-SIG | pose_source 0–3 confidence | All |
| 09 | Rights | RT-09 · AF-SIG | Consent scope, third-party transfer | All |
PS-08 pose source and RT-09 rights are recorded with the same standing as every other signal. Which values were measured and which were estimated, and what the data may be used for, travel inside the data itself.
We write down what we learned, as it was
Capture protocols, rig comparisons, schema change logs. We write up what did not work too.
- The aeffer-v1 schema — why pose_source and the /derived/ split existMix measured and estimated values in one field and the buyer cannot verify quality. We separate provenance into its own field.
- Capture rigs, three tiers compared — mocopi · HaritoraX · Rokoko · Xsens · Vision ProT2 is the minimum configuration that can produce wrist SE(3). T1 is for practising consent and rights, and falls short of sellable quality.
- Why glasses alone cannot be sold — hand pose quality thresholds and monocular depth errorMonocular depth error of 5–11cm. Real-robot success drops from 95% on stereo to 45% on monocular.
- Why DRM is impossible on training data, and what to do insteadData absorbed into the weights cannot be recalled. We go with three things: compute-to-data, contracts, and a per-customer fingerprint.
Planned datasets
No datasets have been published yet.
Aeffer begins its first pilot capture in Q4 2026. The first dataset will be 50 hours of household tasks in Korean living environments, consisting only of data consented for commercial training.
What you can see today
- Synthetic sample viewerSame schema as real captures
- aeffer-v1 schema specificationFull field list published
- Capture roadmapWhich tasks we collect first
PLANNED · ROADMAP
aeffer/ko-logistics-20h
Planned Q1 2027Logistics box handling
aeffer/ko-kitchen-15h
Planned Q1 2027Kitchen movement · simultaneous third-person capture
Why Aeffer
Æffer
What comes in and what goes out
shared root ferre · to carry (Latin)
Æ is a ligature (ash) from Latin and Old English. Just as the two directions share one root, the two letters sit joined in one. We carry from people, and we carry to robots. What sits between the two is us.
I would like to
provide data
You wear a rig and demonstrate everyday tasks. Payment is based on approved hours, you set your own consent scope, and you can withdraw at any time.
- Paid on approved hours · 30-minute sessions
- You choose the consent scope · withdrawable
- Capture equipment lent free of charge
We need
data
Tell us the tasks and signal specs you need and we will design the capture plan with you. Pilots start from small bespoke captures.
- Bespoke capture per task · specs agreed with you
- Licensed for commercial training
- compute-to-data or export
