Right-arm cup grasp, second round. About 25–30 new episodes. Instructions first; the evidence behind them is at the bottom if you want it.
Most existing episodes sit at one end of the table. Put the cup at the thin bands in the map below and record until each has about 12 episodes.
That is roughly +25 to +30 episodes, almost all at the three thin bands. Nothing about how you record changes for this item — just move the cup there.
Right now every episode runs to the same clock, which lets the policy memorise timing instead of watching the cup. Two things, both while recording the episodes from item 1:
Do not change where the arm starts. An earlier version of this page asked for varied starting poses. That was wrong on measurement: the existing episodes already vary a lot at the start (shoulder-lift, elbow and wrist-flex all spread 15–19°), and the one axis that is consistent — which way the arm points left-right — has no evidence behind changing it. Deliberately starting turned away from the cup also asks the demonstrator to perform a motion they have never practised, so the quality of the demonstration itself tends to drop. Start wherever the arm naturally sits.
Approach deliberately off to one side, notice it, correct, then complete the grasp. All 40 existing episodes are clean first-time successes, so the policy has never seen a correction.
Which band, or a table coordinate — one metadata field per episode. Without it, any later analysis has to guess where the cup was from the arm’s own pose, which is circular.
Top-down. Circle area and fill = episodes recorded at that band. A dashed red ring marks a band below the target of 12. Bands are ordered by the arm’s shoulder-pan angle at the moment of grasp.
The dashed line is the target of 12 episodes per band — roughly what the densest band already has.
python grasp_sight_check.py --dataset <name>Eight autonomous trials on 2026-09-08 scored 0/8. Two measurements from the existing 40 episodes account for most of it.
Sixteen episodes sit in the densest band and one in the sparsest — see the map above. The operator called the sparse end a “hard zone”, hard even when recording demonstrations. The measurement says it is not intrinsically hard, it is nearly unrecorded: at that band the policy swung the arm through 1139° of reach (every other band: 341–725°) and never closed the gripper once.
| Measured across all 40 episodes | Median | Spread |
|---|---|---|
| Gripper opens | 5.30 s | ±1.10 s |
| Gripper closes | 8.45 s | ±1.42 s |
| Closes at …% through the episode | 63.2% | ±4.6% |
A 4.6% spread makes “close at 63% through” a near-perfect rule — a policy can learn the clock and drive training loss down without ever locating the cup. Timing correlates with cup position at only r = 0.16–0.26, so it carries almost no information about where the cup is.
On the robot this breaks immediately: demonstrations were recorded at 30 Hz, the control loop runs at about 4 Hz, so a learned tempo arrives at the wrong moment. That is exactly what was observed — on one trial the gripper opened too late and pushed the cup away; on the next it opened and closed too early and stopped short.
In the current dataset the left arm reads 0.000 in all 16&thinspace487 frames, so its recorded standard deviation is exactly zero. Normalisation divides by std + 1e-8, so on the real robot — where the left arm is connected and reports real angles — those readings are amplified roughly 1e8×. Measured: moving the left arm by 0.1° changed the right arm’s commanded output by 64°. It is patched on the inference side, but the next dataset should not reintroduce it.
xlerobot-team/xlerobot-right-pick-cup-20260826-0042 (40 episodes / 16&thinspace487 frames)
and eight trials on 2026-09-08 (trials-20260908-172208.json, 0/8).