The complete chain on this machine, from a camera pixel to a closed gripper, with no policy anywhere in it. Every step names the file that already implements it, or says plainly that it does not exist yet. The second half answers a separate question: what the same geometry buys if SmolVLA is left in the loop for the final grasp. 在我们这台机器上,从相机像素到夹爪闭合的完整链路,中间不放任何策略。 每一步都点名仓库里已经实现它的那个文件,或者直说它还不存在。 后半页回答另一个问题:如果最后的抓取仍然交给 SmolVLA,同一套几何能买到什么。
Six of the eight steps exist and are tested. The two that are missing are both
geometry, not code volume: the transform from camera frame to arm base (hand-eye, never measured) and
lateral solving (shoulder_pan — there is no atan2 anywhere in the repository).
The intrinsics question has a trap in it: the head colour camera was calibrated on 2026-09-04, the depth camera was not,
and the colour numbers must not be pasted into the depth path.
八步里有六步已经存在并且验证过。缺的两步都是几何,不是代码量:
相机系到臂基座的变换(手眼,从没量过),以及横向求解(shoulder_pan —— 整个仓库里没有一处 atan2)。
内参这条里有个坑:头部彩色相机 2026-09-04 标定完了,深度相机没有,而彩色那组数不能填进深度链路。
Cup on the table → gripper closed on it. No SmolVLA, no learned component of any kind after the detector. 杯子在桌上 → 夹爪合在它上面。检测器之后没有 SmolVLA,也没有任何学出来的部件。
trial preset把头移到 trial 预置位
The camera is mounted on a moving head, so any camera-to-base transform is only valid at one head pose.
Every calibration and every run has to start from the same preset. The preset has not been saved yet — one minute on site.
Note that run_policy_trials.py's head_lock reads wherever the head happens to be at startup; it is not a fixed value.
相机装在会动的头上,所以任何相机→基座的变换只在某一个头部姿态下成立。
每次标定和每次运行都必须从同一个预置位开始。这个预置位还没存——现场一分钟的事。
注意 run_policy_trials.py 的 head_lock 是启动时读当前位置就锁在哪,不是固定值。"a cup". Use OWLv2 (google/owlv2-base-patch16-ensemble) at threshold 0.30.
Measured on 36 of our own frames, every box checked by eye after cropping: 19 boxes, all 19 are cups (100% precision), 16 of them the target cup.
The VR controller and the table grommet that wrecked OWL-ViT and Grounding DINO produce zero boxes under OWLv2.
Detection rate per frame: 12/12 on the 0826 footage (the batch the deployed model was trained on) and 9/12 on 0903.
"a cup"。用 OWLv2(google/owlv2-base-patch16-ensemble),阈值 0.30。
在我们自己的 36 帧上实测、每个框裁出来放大人工核对过:19 个框,19 个全是杯子(精度 100%),其中 16 个是目标杯。
把 OWL-ViT 和 Grounding DINO 毁掉的 VR 手柄和桌面穿线圆孔,在 OWLv2 下一个框都不出。
每帧检出率:0826 那批画面(现役模型就是用它训的)12/12,0903 那批 9/12。
X = (u − cx)·z / fx, Y = (v − cy)·z / fy. Three lines of arithmetic — the difficulty is entirely in the four numbers.
Read this before you fill them in. The head colour camera was calibrated on 2026-09-04:
fx 459.59, fy 459.79, cx 323.85, cy 240.17, radtan distortion, RMS reprojection 0.335 px over 17 images
(05-training/onsite/head_intrinsics-2026-09-04.json). Those numbers do not apply to the depth stream.
On the Gemini 335 the depth and the colour images come from different sensors with different fields of view, and
cup_locate_demo.py:36-47 says so explicitly: pasting the colour intrinsics into the depth path makes the
coordinates more wrong, just less obviously so. What is in the code today is FX = FY = 320.0, an
assumed ~90° horizontal field of view — a guess. So Z is trustworthy and X/Y are approximate.
X = (u − cx)·z / fx,Y = (v − cy)·z / fy。三行算术——难的全在那四个数上。
填之前先看这段。头部彩色相机 2026-09-04 标定完了:
fx 459.59, fy 459.79, cx 323.85, cy 240.17,radtan 畸变,17 张图 RMS 重投影 0.335 px
(05-training/onsite/head_intrinsics-2026-09-04.json)。这组数不适用于深度流。
Gemini 335 上深度和彩色来自不同的传感器、视场也不同,cup_locate_demo.py:36-47 写得很明确:
把彩色内参填到深度链路里会让坐标错得更离谱,只是错得没那么明显。
代码里现在是 FX = FY = 320.0,按水平视角约 90° 假设出来的——是猜的。所以 Z 可信,X/Y 只是近似。
trial head pose.
not measured
这是手眼变换,从来没量过。它是这条链路里最大的一个拦路项。
两条路:经典的棋盘格手眼(棋盘夹在爪子上,不是固定在桌上——09-04 给的指示把这个说反了),
或者从已录的 202 集里离线拟合,不占现场时间。
记住第 1 步:拟合出来的东西只在 trial 那个头部姿态下成立。
未量03-software/scripts/arm_kinematics.py, and do not import the upstream kinematics directly.
XLeRobot's SO101Kinematics.forward_kinematics and inverse_kinematics are not inverses of each other:
measured reproducibly across the whole workspace, IK(FK(j2, j3)) → (j2, −j3 + 32.35°). Feeding a pose through both lands
up to 196 mm away from where it started. The hardware follows inverse_kinematics, so IK defines the convention and
it is FK that is wrong. arm_kinematics.forward() corrects it; selftest() round-trips the whole reachable workspace.
Link lengths come from the upstream model, full extension 0.2509 m.
inverse(y, z) → (shoulder_lift, elbow_flex), where y is forward reach and z is height relative to the
shoulder pivot. It does not solve shoulder_pan, and it does not choose the wrist angles.
There is no atan2 anywhere in this repository — nothing currently turns a lateral offset into a pan angle.
You will have to add: pan = atan2(x, y) about the shoulder pivot (the pivot's offset from the base centre line is
also unmeasured), and a choice of wrist_flex that gives the gripper the pitch you want at the cup.
planar part measured and self-tested pan and wrist: nothing exists
用 03-software/scripts/arm_kinematics.py,不要直接 import 上游的运动学。
XLeRobot 的 SO101Kinematics.forward_kinematics 和 inverse_kinematics 互相不是逆:
在整个工作空间里可复现地实测到 IK(FK(j2, j3)) → (j2, −j3 + 32.35°)。一个位姿过一遍这两个函数,
落点最远偏 196 mm。硬件跟的是 inverse_kinematics,所以 IK 定义了约定,错的是 FK。
arm_kinematics.forward() 把它纠正过来;selftest() 对整个可达工作空间做往返验证。
连杆长度取自上游模型,完全伸展 0.2509 m。
inverse(y, z) → (shoulder_lift, elbow_flex),其中 y 是前伸距离,z 是相对肩关节枢轴的高度。
它不解 shoulder_pan,也不选手腕角度。
整个仓库里没有一处 atan2——目前没有任何代码把横向偏移换算成 pan 角。
你要自己加:绕肩枢轴的 pan = atan2(x, y)(枢轴相对底盘中线的偏移同样没量过),
以及一个能让爪子在杯子处呈你想要的俯仰角的 wrist_flex。
平面那部分已实测并自测 pan 和手腕:什么都没有safe_replay.py's Robot.send(targets) wraps XLerobot.send_action and is the primitive to copy.
detect_ports() identifies the two buses by servo count, not by ACM number, which is the only reliable way here.
Action vector order is policy_safety.DIMS (17-D, left arm 0–5, right arm 6–11, head 12–13, base 14–16).
Units trap: the grippers are RANGE_0_100, i.e. percent of travel, not degrees; every other joint is degrees.
An earlier version of the safety code ran the degree formula on the grippers and silently deleted a third of their travel.
One program per bus — two programs on one /dev/ttyACM* corrupt each other's packets and it shows up as phantom hardware faults.
in use
safe_replay.py 里的 Robot.send(targets) 包了 XLerobot.send_action,照抄这个原语就行。
detect_ports() 按舵机数认总线,不认 ACM 编号——在这台机器上这是唯一可靠的办法。
动作向量顺序见 policy_safety.DIMS(17 维,左臂 0–5,右臂 6–11,头 12–13,底盘 14–16)。
单位陷阱:夹爪是 RANGE_0_100,也就是行程百分比,不是角度;其余关节都是度。
安全层早期版本对夹爪套了角度公式,无声地删掉了三分之一的行程。
一条总线一个程序——两个程序同时开一个 /dev/ttyACM* 会互相打乱包,表现成莫名其妙的硬件故障。
在用policy_safety.PolicySafetyFilter. It was written for policy output, but a scripted IK motion should go through it too —
a wrong sign in your transform is exactly the failure it catches. It does: NaN/inf/absurd-magnitude frame rejection, base
zeroing, per-joint step cap, and a Cartesian envelope checked through FK — out of bounds means reject the frame and hold,
not clamp. The floor is computed from the gripper's pitch, pitch = wrist_flex + shoulder_lift + elbow_flex, giving
z_floor(pitch) and y_floor(pitch_y). That detail matters: an earlier version tested a horizontal gripper against a
vertical-gripper floor and rejected 96.6% of frames; with the pitch-aware floor it is 0.49%.
Envelope values live in configs/teleop_safety.yaml.
in use
policy_safety.PolicySafetyFilter。它是为策略输出写的,但脚本化的 IK 动作同样应该过它——
你的变换里符号写反,正好就是它能抓住的那类错。它做:NaN/inf/绝荒量级拒帧、底盘置零、逐关节步长上限,
以及通过 FK 检查笛卡尔包络——越界是拒帧并保持,不是钳位。
地板按爪子的俯仰角算,pitch = wrist_flex + shoulder_lift + elbow_flex,给出
z_floor(pitch) 和 y_floor(pitch_y)。这个细节很关键:早期版本拿「爪子竖直」的地板去卡水平的爪子,
拒了 96.6% 的帧;换成按俯仰角算之后是 0.49%。
包络的数值在 configs/teleop_safety.yaml。
在用| Item事项 | Status状态 | Where / what it costs在哪 / 要花什么 |
|---|---|---|
| Head colour camera intrinsics头部彩色相机内参 | ✅ done已完成 | head_intrinsics-2026-09-04.json · fx 459.59 · RMS 0.335 px · 17 images 张图 |
| Head depth camera intrinsics头部深度相机内参 | ❌ not done没做 | Do not paste the colour numbers in. SDK intrinsics or depth-to-colour alignment, else ~30 min with a checkerboard别把彩色那组填进去。要 SDK 内参或深度对齐彩色,否则棋盘格约 30 分钟 |
| Hand-eye (camera → arm base)手眼(相机 → 臂基座) | ❌ never measured从没量过 | Board clamped in the gripper; or fit offline from the 202 episodes (no site time)棋盘夹在爪子上;或从 202 集离线拟合(不占现场时间) |
Head trial preset头部 trial 预置位 |
❌ not saved没存 | 1 minute on site — but the extrinsic is meaningless without it现场 1 分钟——但没有它外参就没有意义 |
| Depth ↔ RGB time alignment深度 ↔ RGB 时间对齐 | ❌ never verified从没验过 | Wave a hand in front of the camera, compare the two streams在相机前挥手,比对两路流 |
| Planar IK (lift + elbow)平面 IK(lift + elbow) | ✅ works, self-tested可用,有自测 | arm_kinematics.py · fixes a 196 mm upstream bug修掉了上游一个 196 mm 的坑 |
Lateral solving (shoulder_pan)横向求解(shoulder_pan) |
❌ nothing in the repo仓库里没有 | No atan2 anywhere. Shoulder pivot offset also unmeasured一处 atan2 都没有。肩枢轴偏移也没量 |
| Wrist orientation choice手腕姿态的选择 | ❌ nothing没有 | Sets the gripper pitch, which also sets the safety floor它决定爪子俯仰角,而俯仰角决定安全地板 |
| Joint command primitive关节命令原语 | ✅ in use在用 | safe_replay.Robot.send → XLerobot.send_action |
| Safety layer安全层 | ✅ in use在用 | policy_safety.PolicySafetyFilter · reject-and-hold, pitch-aware floor拒帧保持,地板按俯仰角算 |
| Cup detection杯子检测 | ✅ measured已实测 | OWLv2 @0.30 · 100% precision on 36 hand-checked framesOWLv2 @0.30 · 36 帧人工核对,精度 100% |
| Depth capture深度采集 | ✅ works可用 | 79–84% valid pixels after moving to USB3换 USB3 之后有效像素 79–84% |
| Grasp sequence (open/descend/close/lift)抓取序列(张开/下降/闭合/提起) | ❌ to write待写 | Small, no unknowns — but nothing exists不大,也没有未知数——但确实不存在 |
The two offline items unblock the most and cost no site time: fit the hand-eye transform from the 202 recorded
episodes, and add pan solving. The depth intrinsics may turn out to be free — check whether the SDK exposes them or whether
depth-to-colour alignment can simply be turned on, before booking half an hour with a checkerboard.
Saving the trial head preset takes a minute but everything else depends on it, so it should be the first thing done on site.
两件纯离线的事解锁得最多、且不占现场时间:从已录的 202 集里拟合手眼变换,以及补上 pan 求解。
深度内参有可能是免费的——在预约半小时拍棋盘之前,先查 SDK 有没有直接给,或者能不能直接打开深度对齐彩色。
存 trial 头部预置位只要一分钟,但其它一切都依赖它,所以现场第一件事就该做它。
Same geometry as above, but stopping at step 6: drive the arm to a pre-grasp pose over the cup, then let the trained policy do the descend, close and lift. The question worth asking about this variant is narrow: does it make the grasp more reliable inside the range the policy was trained on? 和上面同一套几何,但停在第 6 步:把手臂开到杯子上方的预抓取位,然后让训练好的策略去做下降、闭合、提起。 关于这个变体值得问的问题很窄:在策略训练过的范围内,它能不能让抓取更有把握?
It shortens the part the policy has to get right. Today the policy runs the whole trajectory from the home pose; handing it a pose already above the cup removes the approach phase, which is where the variance accumulates. 它缩短了策略必须做对的那一段。现在策略要从起始位跑完整条轨迹; 把它接手的位置放在杯子正上方,approach 那一段就被拿掉了——而方差正是在那一段累积的。
There is a measurement that points the same way. In the 40 right-arm demonstrations the gripper opens and closes exactly once, in 100% of episodes: open → close → lift. In the 2026-09-04 trials, half the runs showed 2 or 3 open/close cycles — the policy was searching, not executing. 有一个测量指向同一个方向。在右臂那 40 集示范里,夹爪恰好开合一次,100% 的集都是: 张开 → 闭合一次 → 提起。而 2026-09-04 的试跑里,有一半出现了 2 到 3 次开合——策略在试探,不是在执行。
But that evidence is confounded and must not be leaned on. Those same trials sent the task string
"…with the left arm" to a model whose training set contains only "…with the right arm", and ran at speed 0.3,
which the traces show clipped 11–20% of the commanded shoulder-lift motion. Either of those alone could produce the searching
behaviour. The correct rerun has not happened yet.
但这条证据是被污染的,不能靠它。那几次试跑下发的指令串是 "…with the left arm",
而模型的训练集里只有 "…with the right arm";speed 用的 0.3,轨迹实测把要求的下降动作削掉了 11–20%。
这两条里任何一条单独都足以产生那种试探行为。正确的重跑还没做。
Every demonstration starts from the home pose and runs the full trajectory. Handing the policy a mid-trajectory state at t = 0 is out of distribution in time, even when every joint value is inside the trained range. Nothing in this project has tested that. It is a five-minute experiment: place the arm at a pre-grasp pose by hand, start the policy, see whether it picks up the thread or flails. Do that before building anything on this variant. untested 每一集示范都从起始位开始、跑完整条轨迹。把一个轨迹中段的状态当作 t = 0 交给策略, 在时间这个维度上就是分布外的,哪怕每个关节值都落在训练范围内。这个项目里没有任何东西测过这一点。 这是个五分钟的实验:手动把手臂摆到预抓取位,启动策略,看它是接上了还是乱抓。 在这个变体上搭任何东西之前,先做这一步。未测
Only precision. After the handover it is still SmolVLA driving the gripper, so a cup position the policy has
never seen stays out of reach no matter how precisely the arm was placed. The right arm's demonstrations cover only 35.1° of
lateral shoulder_pan (std 5.15°), measured at the frame where the gripper closes. Moving the arm to a cup outside that
band puts the handover state outside the training distribution in exactly the dimension the route was supposed to generalise over.
只有精度。交棒之后握着夹爪的仍然是 SmolVLA,所以策略没见过的杯位,无论手臂摆得多准都还是够不着。
右臂示范在夹爪闭合那一帧测出来的横向 shoulder_pan 覆盖只有 35.1°(std 5.15°)。
把手臂开到这个带之外的杯子那里,交棒时的状态就落在训练分布之外——而且正好落在这条路线本来想泛化的那一维上。
So the honest framing is: IK + SmolVLA is a reliability aid inside the trained envelope, not an answer to new positions. For new positions there are two real options and this is not one of them — go fully geometric (sections 1 and 2 above, no policy at all), or record demonstrations across a wider band of cup positions. 所以老实的说法是:IK + SmolVLA 是训练包络之内的可靠性辅助,不是新位置的答案。 新位置只有两个真正的选项,而这不是其中之一——要么走纯几何(上面第 1、2 节,完全不用策略), 要么在更宽的杯位范围里补录示范。
measured实测 run on this machine, raw data on disk本机跑过,有原始数据 not done未做 stated from the code or a document, not verified by us照代码或文档写的,我们自己没验过 inferred推断 reasoning, not in any document推断,文档里没有
Sources:来源:
03-software/scripts/arm_kinematics.py · cup_locate_demo.py · safe_replay.py ·
policy_safety.py · configs/teleop_safety.yaml ·
05-training/onsite/head_intrinsics-2026-09-04.json ·
05-training/detect/owlv2/records.json · 05-training/detect/manual_review.json ·
GOAL-DEPTH-INTEGRATION.md · GOAL-B0-OBJECT-TOKENS.md · GOAL-T1-DETECTOR-LATENCY.md