← reports index报告首页
2026-09-05 · for Ville · depth line给 Ville · 深度这条线

Picking a cup with IK alone只用 IK 抓放杯子

The complete chain on this machine, from a camera pixel to a closed gripper, with no policy anywhere in it. Every step names the file that already implements it, or says plainly that it does not exist yet. The second half answers a separate question: what the same geometry buys if SmolVLA is left in the loop for the final grasp. 在我们这台机器上,从相机像素到夹爪闭合的完整链路,中间不放任何策略。 每一步都点名仓库里已经实现它的那个文件,或者直说它还不存在。 后半页回答另一个问题:如果最后的抓取仍然交给 SmolVLA,同一套几何能买到什么。

Short version一句话

Six of the eight steps exist and are tested. The two that are missing are both geometry, not code volume: the transform from camera frame to arm base (hand-eye, never measured) and lateral solving (shoulder_pan — there is no atan2 anywhere in the repository). The intrinsics question has a trap in it: the head colour camera was calibrated on 2026-09-04, the depth camera was not, and the colour numbers must not be pasted into the depth path. 八步里有六步已经存在并且验证过。缺的两步都是几何,不是代码量: 相机系到臂基座的变换(手眼,从没量过),以及横向求解(shoulder_pan —— 整个仓库里没有一处 atan2)。 内参这条里有个坑:头部彩色相机 2026-09-04 标定完了,深度相机没有,而彩色那组数不能填进深度链路。

1. The chain, step by step链路逐步拆开

Cup on the table → gripper closed on it. No SmolVLA, no learned component of any kind after the detector. 杯子在桌上 → 夹爪合在它上面。检测器之后没有 SmolVLA,也没有任何学出来的部件。

  1. Put the head in the trial preset把头移到 trial 预置位 The camera is mounted on a moving head, so any camera-to-base transform is only valid at one head pose. Every calibration and every run has to start from the same preset. The preset has not been saved yet — one minute on site. Note that run_policy_trials.py's head_lock reads wherever the head happens to be at startup; it is not a fixed value. 相机装在会动的头上,所以任何相机→基座的变换只在某一个头部姿态下成立。 每次标定和每次运行都必须从同一个预置位开始。这个预置位还没存——现场一分钟的事。 注意 run_policy_trials.py 的 head_lock 是启动时读当前位置就锁在哪,不是固定值。
  2. Find the cup in the head RGB image在头部 RGB 图上找到杯子 Open-vocabulary detector, text prompt "a cup". Use OWLv2 (google/owlv2-base-patch16-ensemble) at threshold 0.30. Measured on 36 of our own frames, every box checked by eye after cropping: 19 boxes, all 19 are cups (100% precision), 16 of them the target cup. The VR controller and the table grommet that wrecked OWL-ViT and Grounding DINO produce zero boxes under OWLv2. Detection rate per frame: 12/12 on the 0826 footage (the batch the deployed model was trained on) and 9/12 on 0903.
    1623 ms per frame, and that is fine here. In this chain the detector runs once, before the arm starts moving: look, fix where the cup is, then drive to that coordinate. The cup does not move during the grasp, so there is no need to keep looking. Waiting a second and a half before the motion starts costs nothing.
    What is not fine is running it every frame. The robot produces one chunk of motion every 1667 ms and SmolVLA's own inference already takes 1329 ms of that, leaving 337 ms. A 1623 ms detector blows that budget by nearly 5×, so the arm would sit waiting on the detector and the motion would come out in stutters. Look-once-then-move works on this board; keep-looking does not. measured
    开放词表检测器,文本 prompt "a cup"。用 OWLv2(google/owlv2-base-patch16-ensemble),阈值 0.30。 在我们自己的 36 帧上实测、每个框裁出来放大人工核对过:19 个框,19 个全是杯子(精度 100%),其中 16 个是目标杯。 把 OWL-ViT 和 Grounding DINO 毁掉的 VR 手柄和桌面穿线圆孔,在 OWLv2 下一个框都不出。 每帧检出率:0826 那批画面(现役模型就是用它训的)12/12,0903 那批 9/12。
    单帧 1623 ms,在这里没问题。这条链路里检测只在手臂开始动之前跑一次: 看一眼,定下杯子在哪,然后照这个坐标去。抓的过程中杯子不动,所以不需要一直看。 动作开始前等一秒半,不花什么代价。
    不行的是每帧都跑。机器人每 1667 ms 产出一段动作,其中 SmolVLA 自己推理就用掉 1329 ms,只剩 337 ms。 1623 ms 的检测器把这个预算撑爆将近 5 倍——手臂会一直等检测器,动作就断断续续。 「看一眼再动」在这块板子上成立,「一直看」不成立。 实测
  3. Read the depth under those pixels读这些像素下面的深度 640×480 16-bit PNG, millimetres. Clip outliers to 100–2000 mm before taking a median. Depth capture itself works — 79–84% valid pixels after the move to USB3. Either take the median inside the box, or prompt SAM with the box first and take the median inside the mask, which is cleaner but adds a second model. 640×480 16-bit PNG,单位毫米。取中位数之前先按 100–2000 mm 裁掉离群值。 深度采集本身是通的——换到 USB3 之后有效像素 79–84%。要么在框内取中位数, 要么先用框提示 SAM 取掩码内的中位数,后者更干净但多一个模型。
  4. Back-project to a 3D point in the camera frame反投影成相机系的 3D 点 X = (u − cx)·z / fx, Y = (v − cy)·z / fy. Three lines of arithmetic — the difficulty is entirely in the four numbers. Read this before you fill them in. The head colour camera was calibrated on 2026-09-04: fx 459.59, fy 459.79, cx 323.85, cy 240.17, radtan distortion, RMS reprojection 0.335 px over 17 images (05-training/onsite/head_intrinsics-2026-09-04.json). Those numbers do not apply to the depth stream. On the Gemini 335 the depth and the colour images come from different sensors with different fields of view, and cup_locate_demo.py:36-47 says so explicitly: pasting the colour intrinsics into the depth path makes the coordinates more wrong, just less obviously so. What is in the code today is FX = FY = 320.0, an assumed ~90° horizontal field of view — a guess. So Z is trustworthy and X/Y are approximate.
    Two ways out, either is fine: (a) ask the SDK for the depth stream's own intrinsics, or turn on depth-to-colour alignment and then the calibrated colour numbers do apply; (b) calibrate the IR/depth stream against a checkerboard the same way the colour one was done on 09-04, about 30 minutes on site. not done
    X = (u − cx)·z / fx,Y = (v − cy)·z / fy。三行算术——难的全在那四个数上。 填之前先看这段。头部彩色相机 2026-09-04 标定完了: fx 459.59, fy 459.79, cx 323.85, cy 240.17,radtan 畸变,17 张图 RMS 重投影 0.335 px (05-training/onsite/head_intrinsics-2026-09-04.json)。这组数不适用于深度流。 Gemini 335 上深度和彩色来自不同的传感器、视场也不同,cup_locate_demo.py:36-47 写得很明确: 把彩色内参填到深度链路里会让坐标错得更离谱,只是错得没那么明显。 代码里现在是 FX = FY = 320.0,按水平视角约 90° 假设出来的——是猜的。所以 Z 可信,X/Y 只是近似。
    两条出路,哪条都行:(a) 向 SDK 要深度流自己的内参,或者打开深度对齐到彩色,那样标好的彩色内参就适用了; (b) 像 09-04 标彩色那样,用棋盘格对 IR/深度流单独标一次,现场约 30 分钟。 未做
  5. Transform camera frame → arm base frame相机系 → 臂基座系的变换 This is the hand-eye transform and it has never been measured. It is the single largest blocker in this chain. Two routes: the classic checkerboard hand-eye (the board is clamped in the gripper, not fixed to the table — the instruction given on 09-04 had this backwards), or fit it offline from the 202 recorded episodes, which needs no site time. Remember step 1: whatever you fit is only valid at the trial head pose. not measured 这是手眼变换,从来没量过。它是这条链路里最大的一个拦路项。 两条路:经典的棋盘格手眼(棋盘夹在爪子上,不是固定在桌上——09-04 给的指示把这个说反了), 或者从已录的 202 集里离线拟合,不占现场时间。 记住第 1 步:拟合出来的东西只在 trial 那个头部姿态下成立。 未量
  6. 3D point → joint angles (this is the IK)3D 点 → 关节角(这就是 IK) Use 03-software/scripts/arm_kinematics.py, and do not import the upstream kinematics directly. XLeRobot's SO101Kinematics.forward_kinematics and inverse_kinematics are not inverses of each other: measured reproducibly across the whole workspace, IK(FK(j2, j3)) → (j2, −j3 + 32.35°). Feeding a pose through both lands up to 196 mm away from where it started. The hardware follows inverse_kinematics, so IK defines the convention and it is FK that is wrong. arm_kinematics.forward() corrects it; selftest() round-trips the whole reachable workspace. Link lengths come from the upstream model, full extension 0.2509 m.
    The limitation you need to know about: this is 2-DOF planar IK only. inverse(y, z) → (shoulder_lift, elbow_flex), where y is forward reach and z is height relative to the shoulder pivot. It does not solve shoulder_pan, and it does not choose the wrist angles. There is no atan2 anywhere in this repository — nothing currently turns a lateral offset into a pan angle. You will have to add: pan = atan2(x, y) about the shoulder pivot (the pivot's offset from the base centre line is also unmeasured), and a choice of wrist_flex that gives the gripper the pitch you want at the cup. planar part measured and self-tested pan and wrist: nothing exists
    用 03-software/scripts/arm_kinematics.py,不要直接 import 上游的运动学。 XLeRobot 的 SO101Kinematics.forward_kinematics 和 inverse_kinematics 互相不是逆: 在整个工作空间里可复现地实测到 IK(FK(j2, j3)) → (j2, −j3 + 32.35°)。一个位姿过一遍这两个函数, 落点最远偏 196 mm。硬件跟的是 inverse_kinematics,所以 IK 定义了约定,错的是 FK。 arm_kinematics.forward() 把它纠正过来;selftest() 对整个可达工作空间做往返验证。 连杆长度取自上游模型,完全伸展 0.2509 m。
    你必须知道的限制:这是二自由度平面 IK。 inverse(y, z) → (shoulder_lift, elbow_flex),其中 y 是前伸距离,z 是相对肩关节枢轴的高度。 它不解 shoulder_pan,也不选手腕角度。 整个仓库里没有一处 atan2——目前没有任何代码把横向偏移换算成 pan 角。 你要自己加:绕肩枢轴的 pan = atan2(x, y)(枢轴相对底盘中线的偏移同样没量过), 以及一个能让爪子在杯子处呈你想要的俯仰角的 wrist_flex。 平面那部分已实测并自测 pan 和手腕:什么都没有
  7. Send the joint targets下发关节目标 safe_replay.py's Robot.send(targets) wraps XLerobot.send_action and is the primitive to copy. detect_ports() identifies the two buses by servo count, not by ACM number, which is the only reliable way here. Action vector order is policy_safety.DIMS (17-D, left arm 0–5, right arm 6–11, head 12–13, base 14–16). Units trap: the grippers are RANGE_0_100, i.e. percent of travel, not degrees; every other joint is degrees. An earlier version of the safety code ran the degree formula on the grippers and silently deleted a third of their travel. One program per bus — two programs on one /dev/ttyACM* corrupt each other's packets and it shows up as phantom hardware faults. in use safe_replay.py 里的 Robot.send(targets) 包了 XLerobot.send_action,照抄这个原语就行。 detect_ports() 按舵机数认总线,不认 ACM 编号——在这台机器上这是唯一可靠的办法。 动作向量顺序见 policy_safety.DIMS(17 维,左臂 0–5,右臂 6–11,头 12–13,底盘 14–16)。 单位陷阱:夹爪是 RANGE_0_100,也就是行程百分比,不是角度;其余关节都是度。 安全层早期版本对夹爪套了角度公式,无声地删掉了三分之一的行程。 一条总线一个程序——两个程序同时开一个 /dev/ttyACM* 会互相打乱包,表现成莫名其妙的硬件故障。 在用
  8. Put every frame through the safety layer每一帧都过安全层 policy_safety.PolicySafetyFilter. It was written for policy output, but a scripted IK motion should go through it too — a wrong sign in your transform is exactly the failure it catches. It does: NaN/inf/absurd-magnitude frame rejection, base zeroing, per-joint step cap, and a Cartesian envelope checked through FK — out of bounds means reject the frame and hold, not clamp. The floor is computed from the gripper's pitch, pitch = wrist_flex + shoulder_lift + elbow_flex, giving z_floor(pitch) and y_floor(pitch_y). That detail matters: an earlier version tested a horizontal gripper against a vertical-gripper floor and rejected 96.6% of frames; with the pitch-aware floor it is 0.49%. Envelope values live in configs/teleop_safety.yaml. in use policy_safety.PolicySafetyFilter。它是为策略输出写的,但脚本化的 IK 动作同样应该过它—— 你的变换里符号写反,正好就是它能抓住的那类错。它做:NaN/inf/绝荒量级拒帧、底盘置零、逐关节步长上限, 以及通过 FK 检查笛卡尔包络——越界是拒帧并保持,不是钳位。 地板按爪子的俯仰角算,pitch = wrist_flex + shoulder_lift + elbow_flex,给出 z_floor(pitch) 和 y_floor(pitch_y)。这个细节很关键:早期版本拿「爪子竖直」的地板去卡水平的爪子, 拒了 96.6% 的帧;换成按俯仰角算之后是 0.49%。 包络的数值在 configs/teleop_safety.yaml。 在用
  9. The grasp sequence itself抓取序列本身 Open the gripper → move to a pre-grasp pose above the cup → descend → close → lift. This does not exist. Everything above gets the arm to a point; nothing in the repository turns that point into a grasp. It is the smallest of the missing pieces and the one with no unknowns in it — but it does have to be written, and the descend and close phases are where the accuracy of steps 4 and 5 actually gets tested. to write 张开夹爪 → 移到杯子上方的预抓取位 → 下降 → 闭合 → 提起。 这段不存在。上面所有步骤只是把手臂送到一个点;仓库里没有任何东西把那个点变成一次抓取。 它是缺的几块里最小的一块,也是唯一没有未知数的一块——但确实要写, 而且第 4、5 步的精度到底够不够,正是在下降和闭合这两个阶段才被真正检验。 待写

2. What is done and what is not哪些做完了、哪些没有

Item事项 Status状态 Where / what it costs在哪 / 要花什么
Head colour camera intrinsics头部彩色相机内参 ✅ done已完成 head_intrinsics-2026-09-04.json · fx 459.59 · RMS 0.335 px · 17 images 张图
Head depth camera intrinsics头部深度相机内参 ❌ not done没做 Do not paste the colour numbers in. SDK intrinsics or depth-to-colour alignment, else ~30 min with a checkerboard别把彩色那组填进去。要 SDK 内参或深度对齐彩色,否则棋盘格约 30 分钟
Hand-eye (camera → arm base)手眼(相机 → 臂基座) ❌ never measured从没量过 Board clamped in the gripper; or fit offline from the 202 episodes (no site time)棋盘夹在爪子上;或从 202 集离线拟合(不占现场时间)
Head trial preset头部 trial 预置位 ❌ not saved没存 1 minute on site — but the extrinsic is meaningless without it现场 1 分钟——但没有它外参就没有意义
Depth ↔ RGB time alignment深度 ↔ RGB 时间对齐 ❌ never verified从没验过 Wave a hand in front of the camera, compare the two streams在相机前挥手,比对两路流
Planar IK (lift + elbow)平面 IK(lift + elbow) ✅ works, self-tested可用,有自测 arm_kinematics.py · fixes a 196 mm upstream bug修掉了上游一个 196 mm 的坑
Lateral solving (shoulder_pan)横向求解(shoulder_pan) ❌ nothing in the repo仓库里没有 No atan2 anywhere. Shoulder pivot offset also unmeasured一处 atan2 都没有。肩枢轴偏移也没量
Wrist orientation choice手腕姿态的选择 ❌ nothing没有 Sets the gripper pitch, which also sets the safety floor它决定爪子俯仰角,而俯仰角决定安全地板
Joint command primitive关节命令原语 ✅ in use在用 safe_replay.Robot.send → XLerobot.send_action
Safety layer安全层 ✅ in use在用 policy_safety.PolicySafetyFilter · reject-and-hold, pitch-aware floor拒帧保持,地板按俯仰角算
Cup detection杯子检测 ✅ measured已实测 OWLv2 @0.30 · 100% precision on 36 hand-checked framesOWLv2 @0.30 · 36 帧人工核对,精度 100%
Depth capture深度采集 ✅ works可用 79–84% valid pixels after moving to USB3换 USB3 之后有效像素 79–84%
Grasp sequence (open/descend/close/lift)抓取序列(张开/下降/闭合/提起) ❌ to write待写 Small, no unknowns — but nothing exists不大,也没有未知数——但确实不存在
Order of work, if it helps如果有用,一个建议的顺序

The two offline items unblock the most and cost no site time: fit the hand-eye transform from the 202 recorded episodes, and add pan solving. The depth intrinsics may turn out to be free — check whether the SDK exposes them or whether depth-to-colour alignment can simply be turned on, before booking half an hour with a checkerboard. Saving the trial head preset takes a minute but everything else depends on it, so it should be the first thing done on site. 两件纯离线的事解锁得最多、且不占现场时间:从已录的 202 集里拟合手眼变换,以及补上 pan 求解。 深度内参有可能是免费的——在预约半小时拍棋盘之前,先查 SDK 有没有直接给,或者能不能直接打开深度对齐彩色。 存 trial 头部预置位只要一分钟,但其它一切都依赖它,所以现场第一件事就该做它。

3. The other option: IK to pre-grasp, then hand over to SmolVLA另一个选项:IK 送到预抓取位,再交棒给 SmolVLA

Same geometry as above, but stopping at step 6: drive the arm to a pre-grasp pose over the cup, then let the trained policy do the descend, close and lift. The question worth asking about this variant is narrow: does it make the grasp more reliable inside the range the policy was trained on? 和上面同一套几何,但停在第 6 步:把手臂开到杯子上方的预抓取位,然后让训练好的策略去做下降、闭合、提起。 关于这个变体值得问的问题很窄:在策略训练过的范围内,它能不能让抓取更有把握?

Reasons it plausibly helps — none of them tested可能有帮助的理由——一条都没测过

It shortens the part the policy has to get right. Today the policy runs the whole trajectory from the home pose; handing it a pose already above the cup removes the approach phase, which is where the variance accumulates. 它缩短了策略必须做对的那一段。现在策略要从起始位跑完整条轨迹; 把它接手的位置放在杯子正上方,approach 那一段就被拿掉了——而方差正是在那一段累积的。

There is a measurement that points the same way. In the 40 right-arm demonstrations the gripper opens and closes exactly once, in 100% of episodes: open → close → lift. In the 2026-09-04 trials, half the runs showed 2 or 3 open/close cycles — the policy was searching, not executing. 有一个测量指向同一个方向。在右臂那 40 集示范里,夹爪恰好开合一次,100% 的集都是: 张开 → 闭合一次 → 提起。而 2026-09-04 的试跑里,有一半出现了 2 到 3 次开合——策略在试探,不是在执行。

But that evidence is confounded and must not be leaned on. Those same trials sent the task string "…with the left arm" to a model whose training set contains only "…with the right arm", and ran at speed 0.3, which the traces show clipped 11–20% of the commanded shoulder-lift motion. Either of those alone could produce the searching behaviour. The correct rerun has not happened yet. 但这条证据是被污染的,不能靠它。那几次试跑下发的指令串是 "…with the left arm", 而模型的训练集里只有 "…with the right arm";speed 用的 0.3,轨迹实测把要求的下降动作削掉了 11–20%。 这两条里任何一条单独都足以产生那种试探行为。正确的重跑还没做。

The reason it might not work at all它可能根本不成立的理由

Every demonstration starts from the home pose and runs the full trajectory. Handing the policy a mid-trajectory state at t = 0 is out of distribution in time, even when every joint value is inside the trained range. Nothing in this project has tested that. It is a five-minute experiment: place the arm at a pre-grasp pose by hand, start the policy, see whether it picks up the thread or flails. Do that before building anything on this variant. untested 每一集示范都从起始位开始、跑完整条轨迹。把一个轨迹中段的状态当作 t = 0 交给策略, 在时间这个维度上就是分布外的,哪怕每个关节值都落在训练范围内。这个项目里没有任何东西测过这一点。 这是个五分钟的实验:手动把手臂摆到预抓取位,启动策略,看它是接上了还是乱抓。 在这个变体上搭任何东西之前,先做这一步。未测

Does it help generalisation, or is it only precision?它帮的是泛化,还是只是精度?

Only precision. After the handover it is still SmolVLA driving the gripper, so a cup position the policy has never seen stays out of reach no matter how precisely the arm was placed. The right arm's demonstrations cover only 35.1° of lateral shoulder_pan (std 5.15°), measured at the frame where the gripper closes. Moving the arm to a cup outside that band puts the handover state outside the training distribution in exactly the dimension the route was supposed to generalise over. 只有精度。交棒之后握着夹爪的仍然是 SmolVLA,所以策略没见过的杯位,无论手臂摆得多准都还是够不着。 右臂示范在夹爪闭合那一帧测出来的横向 shoulder_pan 覆盖只有 35.1°(std 5.15°)。 把手臂开到这个带之外的杯子那里,交棒时的状态就落在训练分布之外——而且正好落在这条路线本来想泛化的那一维上。

So the honest framing is: IK + SmolVLA is a reliability aid inside the trained envelope, not an answer to new positions. For new positions there are two real options and this is not one of them — go fully geometric (sections 1 and 2 above, no policy at all), or record demonstrations across a wider band of cup positions. 所以老实的说法是:IK + SmolVLA 是训练包络之内的可靠性辅助,不是新位置的答案。 新位置只有两个真正的选项,而这不是其中之一——要么走纯几何(上面第 1、2 节,完全不用策略), 要么在更宽的杯位范围里补录示范。

4. Related page相关页面

Two ways to use depth, and how to accept B0深度的两条用法与 B0 验收 The generalisation question in full: whether combining depth geometry with the policy buys generalisation at all (it does not, and why), what pure geometry buys instead, and the separate route of feeding object recognition into the policy as object tokens — which is the one thing that does buy generalisation across different-looking cups. Also carries the acceptance criteria and the A/B design for that route. 泛化那个问题的完整版:把深度几何和策略结合到底买不买得到泛化(买不到,以及为什么)、 纯几何换来的是什么,以及另一条独立的路线——把物体识别作为物体 token 喂进策略, 那是唯一真正能买到「换一只长得不一样的杯子」这种泛化的做法。那一页还带着这条路线的验收判据和 A/B 设计。

5. Evidence levels证据等级

measured实测 run on this machine, raw data on disk本机跑过,有原始数据 not done未做 stated from the code or a document, not verified by us照代码或文档写的,我们自己没验过 inferred推断 reasoning, not in any document推断,文档里没有

Sources:来源: 03-software/scripts/arm_kinematics.py · cup_locate_demo.py · safe_replay.py · policy_safety.py · configs/teleop_safety.yaml · 05-training/onsite/head_intrinsics-2026-09-04.json · 05-training/detect/owlv2/records.json · 05-training/detect/manual_review.json · GOAL-DEPTH-INTEGRATION.md · GOAL-B0-OBJECT-TOKENS.md · GOAL-T1-DETECTOR-LATENCY.md