SmolVLA is the main policy, and generalisation is its biggest potential — but it is not controllable enough yet. Three additions are on the table, and they buy three different things: depth run on its own as a fallback, depth combined with the policy for accuracy, and object recognition to strengthen generalisation. This page says what each one is, what it costs, and — for the object-recognition route — how to actually accept it. SmolVLA 是主策略,泛化是它最大的潜力——但目前还不够可控。 围绕它有三件可以加的事,各买各的:深度单独跑当兜底、深度与策略结合提高准确度、物体识别增强泛化。 这一页说清每件是什么、代价是什么,以及物体识别那条到底怎么验收。
| What you add加什么 | How it is used怎么用 | What it buys买到什么 |
|---|---|---|
| Depth, on its own深度,单独跑 B1 solo |
No policy at all. Look once to locate the cup → IK → a scripted descend, close and lift. 完全不用策略。看一眼定位杯子 → IK → 脚本化的下降、闭合、提起。 | Insurance. If SmolVLA does not work on the day, something on the table can still pick the cup up. 兜底。SmolVLA 当天不 work 的话,台上还有东西能把杯子抓起来。 |
| Depth + SmolVLA深度 + SmolVLA B1 handover |
IK drives the arm to a pre-grasp pose over the cup; the policy does the final grasp. IK 把手臂送到杯子上方的预抓取位,策略负责最后的抓取。 | Accuracy, inside the range the policy was trained on. Not generalisation — see §2. 准确度,在策略训练过的范围之内。不是泛化——见 §2。 |
| Object recognition物体识别 B0 · object tokens |
Cut the picture into objects and give the policy one token per object, alongside the existing image tokens. 把画面按物体切开,每个物体给策略一个 token,和现有的图像 token 并存。 | Generalisation — a cup that looks different from the ones in the training data. 泛化——一只和训练数据里长得不一样的杯子。 |
Generalisation is what SmolVLA can do that geometry cannot — but right now the model learns an association between a patch of the picture and an action: token 137 always means the same square of the image, whatever happens to be sitting there. That is the mechanical reason a different-looking cup breaks it, and it is why generalisation currently feels uncontrollable. Object recognition changes what the model is looking at: it learns "where that object is" instead of "square 137 lit up". That is the optimisation that makes the potential controllable — not a bigger model, and not more compute. 泛化是 SmolVLA 能做而几何做不到的事——但现在模型学的是画面某一块和动作之间的关联: 第 137 号 token 永远代表画面同一个格子,不管那里坐着什么东西。 这就是「换一只杯子就失效」的机械原因,也是泛化现在显得不可控的原因。 物体识别改变的是模型看的东西:它学的变成「那个物体在哪」,而不是「第 137 号格子亮了」。 这才是让那个潜力变得可控的优化——不是更大的模型,也不是更多算力。
No, and no. Both depth uses above read the sensor live, once before the run — nothing goes into a dataset and nothing goes into the network. Object recognition does not touch depth at all; it is pure RGB, needs retraining (~1.3 h / $1) but no new recordings. 都不用。上面两种深度用法都是运行时、开跑前读一次——不入数据集,也不进网络。 物体识别根本不碰深度,是纯 RGB 的;它要重训(约 1.3 小时 / $1),但不用重采数据。
There is a route that feeds depth into training — B2, "Depth Helps" — and it is the one we are not taking, because it requires re-recording every episode in a native depth format. It is archived at the foot of this page with what it is, why it might be worth something, what it would cost and what is still uncertain. 确实有一条把深度喂进训练的路线——B2「Depth Helps」——那条我们不走, 因为它要求把每一集都用原生深度格式重录。它存档在本页页尾:它是什么、为什么可能有价值、要额外做什么、什么不确定。
A purely geometric chain. No change to the policy, no retraining, no new data. Depth is used once before the run, not inside the control loop. 一条纯几何链路,不改策略、不重训、不要新数据。深度只在开跑前用一次,不在控制回路里。
trial preset头移到 trial 预置位
The camera sits on a moving head, so the extrinsic is only valid at one head pose — move the head and it is void. The preset has not been saved yet.
相机装在会动的头上,外参是「头在某个姿态时」的值,头一动全废。这个预置位还没存。"a cup". Use OWLv2 at threshold 0.30 — see the card below for the numbers.
开放词表检测器,prompt "a cup"。用 OWLv2,阈值 0.30——数据见下面那张卡。FX = FY = 320, a guess: Z is trustworthy, X/Y approximate.
要深度相机的内参。彩色相机 2026-09-04 标定完了,但那组数不适用于深度流,不能填进去。
代码里现在是 FX = FY = 320,是猜的:Z 准,X/Y 只是近似。OWLv2 on our own head-camera frames, every box cropped and checked by eye: OWLv2 跑在我们自己的头相机画面上,每个框都裁出来放大人工看过:
| Footage画面 | Cup found找到杯子 |
|---|---|
| 0826 — the batch the deployed model was trained on现役模型就是用这批训的 | 12 / 12 |
| 0903 — sep3sep3 | 9 / 12 |
And nothing that is not a cup gets boxed: across all 36 frames, 19 boxes, 19 of them cups — zero false positives. The VR controller and the table grommet that ruined the two earlier detectors produce no boxes at all under OWLv2. 而且不是杯子的东西不会被框出来:36 帧一共 19 个框,19 个都是杯子——零假阳性。 把之前两个检测器毁掉的 VR 手柄和桌面穿线圆孔,在 OWLv2 下一个框都不出。
Why this detection rate is good enough here. In this route the detector runs once, before the grasp starts: look, fix where the cup is, then the arm goes to that coordinate. The cup does not move during the grasp, so there is no need to keep looking. It takes 1623 ms, and waiting a second and a half before the arm moves costs nothing. 为什么这个检出率在这里够用。这条路线里检测只在开始抓之前跑一次: 看一眼,定下杯子在哪,然后手臂照这个坐标去。抓的过程中杯子不动,所以不需要一直看。 它要 1623 ms,而在手臂动之前等一秒半,不花什么代价。
Why it cannot run continuously. The robot puts out one chunk of motion every 1667 ms, and SmolVLA's own thinking already takes 1329 ms of that — leaving 337 ms. Running a 1623 ms detector every frame overruns that by nearly 5×: the arm would sit waiting on the detector and the motion would come out in stutters. Look once, then move works on this board. Keep looking does not. 为什么不能一直实时检测。机器人每 1667 ms 产出一段动作,其中 SmolVLA 自己思考就用掉 1329 ms——只剩 337 ms。 每帧都跑一个 1623 ms 的检测器,把这个预算超了将近 5 倍:手臂会一直等检测器,动作就断断续续。 看一眼再动在这块板子上成立,一直看不成立。
What we have not measured. Frames from one scene look alike, so a miss tends to be a whole-scene miss. "12 out of 12 frames" is therefore not the same as "finds the cup on every attempt". We only have three scenes — too few to say. 还没测的。同一个场景里的帧长得差不多,所以检不到往往是整场都检不到。 「12 帧里 12 帧」因此不等于「每次抓取都能找到杯子」。我们只有三个场景,太少,说不了。 untested未测
B1's handover version exists to buy new cup positions. But moving the arm to a new position puts
shoulder_pan outside the trained range — the right arm's demonstrations cover only 35.1° of lateral pan (std 5.15°),
measured at the frame where the gripper closes. The state SmolVLA receives at handover therefore lands in exactly the dimension it has never seen.
B1 的交棒版存在的目的是买换位置。可把手臂开到新位置,就意味着 shoulder_pan 被推到训练范围之外——
右臂示范在夹爪闭合那一帧测出的横向覆盖只有 35.1°(std 5.15°)。
SmolVLA 交棒时收到的状态,正好落在它没见过的那一维上。
Operator's ruling, 2026-09-05: this combination is not the answer to new positions. After the handover it is still SmolVLA driving; a position it has never seen stays out of reach however precisely the arm is placed. 操作者 2026-09-05 判定:这个组合不是新位置的答案。交棒之后跑的还是 SmolVLA, 它没见过的位置,手臂摆得再准也够不着。
The five-minute handover test is still worth doing at a cup position inside the trained band — it tests whether the handover mechanism works at all, which matters for the reliability question. What it must not be sold as is a route to position generalisation. That question is treated in full on the IK page, §3. 那个五分钟的交棒测试仍然值得做,但要在训练带之内的杯位上做——它验的是交棒机制本身能不能接上, 这对「更有把握」那个问题有意义。不该被当成通往位置泛化的路子。 这个问题在 IK 那一页 §3 有完整讨论。
The same geometric chain, but carried all the way through — descend, close and lift all driven by IK and a script, never handing over to the policy. Position generalisation is then free: geometry needs no training coverage, and the 35° limit simply does not apply. 同一套几何链路,但一路做到底——下降、闭合、提起全部走 IK + 脚本,中途不交给策略。 这样位置泛化就是 免费的:几何不需要训练数据覆盖,35° 那条限制根本不存在。
What it buys: insurance for the demo. If the policy does not work on the day, there is still something on the table that can pick the cup up. This is a fallback, not a policy improvement — it does not make the model one byte better. 它买的是什么:demo 的保险。策略当天不起效果时,台上还有东西能把杯子抓起来。 这是兜底,不是策略改进——它一个字节也没改善模型。
It costs more than the handover version, not less. With no policy at the end to absorb error, accuracy rests entirely on perception and calibration. None of the geometry prerequisites can be skipped, and the tolerance is tighter. 代价比交棒版更高,不是更低。末端没有策略帮忙兜误差,精度全压在感知和标定上。 几何那几项前置一项都不能省,而且对精度的要求更严。
Detection is no longer the blocker. OWLv2 has zero false positives and finds the cup in every frame of the current footage, and 1623 ms is fine because this route also looks once before moving. The real unknowns are calibration accuracy and the grasp sequence, which has not been written. Full breakdown here. 检测这一关不再是拦路的。OWLv2 零假阳性,在当前画面上每一帧都找到了杯子; 1623 ms 也不碍事,因为这条路同样是动之前看一眼。真正的未知在标定精度,以及那段还没写的抓取序列。 完整拆解在这一页。
Instead of computing a 3D coordinate, cut the picture into objects, give each one a token appended to the prefix, and let the model decide which to look at. No depth, no intrinsics, no hand-eye, no new recordings — but it does need retraining. 不算 3D 坐标,把画面按物体切开,每个物体给一个 token 接到 prefix 末尾,让模型自己挑该看谁。 不用深度、不用内参、不用手眼、不用重采数据,但要重训。
object_proj加一个 object_proj
Copy state_proj (nn.Linear(32, 960), modeling_smolvla.py:508) to project the object features to 960 dims.
照抄 state_proj(nn.Linear(32, 960),modeling_smolvla.py:508),把物体特征投到 960 维。[B, N_obj, 960] at the end of embed_prefix在 embed_prefix 末尾 append [B, N_obj, 960]
modeling_smolvla.py:552 is just a list that gets concatenated at the end. About 15 lines of change.
modeling_smolvla.py:552 那里就是一个 list 最后 concat。改动约 15 行。att_mask must be 1, and the tokens must go lastatt_mask 必须标 1,且必须放在最后
This is the only safe way to write it. att_masks is not an on/off switch for visibility — it is a block boundary marker:
a cumulative sum gives the block number, and token i can see token j iff block(j) ≤ block(i). Marking 1 at the end opens a new block, which
leaves the image and language outputs bit-for-bit unchanged. Inserting before the state token with a 0 instead rewrites the pretrained
visual and language features. verified in the source
这是唯一安全的写法。att_masks 不是「能不能看」的开关,是块边界标记:
cumsum 给出块号,token i 看得到 j ⟺ 块号(j) ≤ 块号(i)。标 1 放最后 = 开新块 = 图像和语言的输出逐比特不变。
插在 state 之前且标 0,会改写预训练的视觉 / 语言特征。读过源码N_obj must be fixedN_obj 必须固定
att_masks is a python list shared across the whole batch. When fewer objects are found, pad with pad_masks=0 —
padding tokens are masked on both the key and the query side, so empty objects contaminate nothing.
att_masks 是 python list,整个 batch 共享同一长度。检出不足时用 pad_masks=0 补——
padding token 在 key 和 query 两侧都被屏蔽,空物体不污染任何东西。object_proj只训 action expert + object_proj
The deployed config is already freeze_vision_encoder: True and train_expert_only: True — SigLIP does not move a parameter,
so this is compatible with the existing training mode. Object tokens are computed offline from the RGB we already recorded, so no re-recording. About 1.3 h / $1.
部署配置本来就是 freeze_vision_encoder: True、train_expert_only: True,SigLIP 一个参数不动,
和现有训练模式兼容。物体 token 从已录的 RGB 离线算,所以不用重采数据。约 1.3 小时 / $1。Speed is not the reason to do this. SmolVLA takes 1329 ms per inference and the three camera streams are only 13% of it (177.7 ms). Cutting tokens saves a fraction of that, and the segmentation model's own cost eats into it — the net could well be negative. The only argument that stands up is generalisation from object-level structure. measured 速度不是理由。SmolVLA 一次推理 1329 ms,三路视觉只占 13%(177.7 ms)。 token 砍下来省的那点,再减去分割模型自己的开销,净收益很可能为负。 唯一站得住的理由是物体级结构先验带来的泛化。实测
Zero-initialisation is not identity. Even with object_proj's last layer zeroed and every value zero, the extra keys still
take part in the softmax normalisation and dilute the existing attention weights. A warmup or a small initial learning rate is still needed.
零初始化不等于恒等。就算 object_proj 最后一层置零、value 全是零,
多出来的 key 仍然参与 softmax 归一化、会稀释原有的注意力权重。仍需 warmup 或小学习率起步。
The B0 document lists work items T1/T2/T3, and T3 says "compute the net benefit and decide" — with no criterion and no controlled design. Here it is. Three hard constraints shape it: B0 那份文档列了 T1 / T2 / T3,T3 写的是「算净收益并决定」——没有判据,也没有对照设计。 下面补上。先说三条决定设计形状的硬约束:
Then three gates, cheapest first. If one fails, stop — do not go on to the next. 然后是三道门,便宜的在前。前一道不过就停,不进下一道。
After training B0, replace every object token with zeros at inference (then repeat with random vectors). 训完 B0 之后,推理时把物体 token 全部换成零向量(再跑一遍随机向量)。
object_proj learned noise. Stop here, no robot needed.
输出几乎不变 → 模型压根没在用这些 token,object_proj 学成了噪声。直接停,不用碰机器人。An open question from the B0 document: how many different cups appear across the existing 202 episodes? If only one, the model has probably locked onto that cup's appearance and B0 has a lot to gain; if several, the model may already have some appearance invariance and there is less to buy. Purely offline — read the keyframes of the merged dataset. B0 文档末尾留的一个问题:现有 202 集里到底用了几种杯子? 全程只有一只 → 模型很可能已经锁死在那只杯子的外观上,B0 的收益空间大; 本来就有好几只 → 模型可能已经有外观不变性,能买到的就少。纯离线——读合并集的关键帧就行。
| Detail内容 | |
|---|---|
| What it is它是什么 | Freeze SmolVLA's SigLIP backbone and add two small modules alongside it: a Depth Completion Module that predicts depth features from RGB via cross-attention, and a Depth-Aware Codebook that discretises depth for noise robustness. Only those two modules are trained. Source: arXiv 2408.05107 把 SmolVLA 的 SigLIP 主干冻住不动,在旁边加两个小模块: Depth Completion Module(用 cross-attention 从 RGB 预测深度特征)和 Depth-Aware Codebook(把深度离散化,抗噪)。 只训这两个小模块。出处 arXiv 2408.05107 |
| Why it might be worth it为什么可能 有价值 | It is the only one of the three that actually feeds depth into the policy. The paper reports 57.95% → 63.95%, with the largest gain on long-horizon tasks (~12%); it needs only about 20 real trajectories per task and we have 30. And at inference it can run on RGB alone (63.15% vs 63.95%, a 0.8-point drop) — depth is only needed during training, so deployment gets simpler, not harder. All three points fit our situation closely. 它是三条里唯一真正把深度喂进策略的。论文报 57.95% → 63.95%,长时程任务提升最大(约 12%); 每个任务只要约 20 条真机轨迹,而我们有 30 集。而且推理时可以只用 RGB(63.15% vs 63.95%,只掉 0.8%)—— 深度只在训练时需要,部署反而更省事。这三点和我们的处境高度吻合。 |
| What extra it would take需要额外 做什么 |
Re-recording. Depth has to go into the dataset as a native feature (orbbec_depth_camera.py already supports it),
and the 30 episodes recorded on 9-02 / 9-03 are still in the old sidecar format and cannot be used. This is the only large cost among the three routes,
and it is exactly why the route was shelved. It also needs three-stage training (warmup / alignment / codebook).
重录数据。深度必须以原生特征录进数据集(orbbec_depth_camera.py 已打通),
而 9-02 / 9-03 那 30 集还是旧的旁路格式,用不了。这是三条路线里唯一一笔大成本,也正是它被搁置的原因。
另外要走三阶段训练(warmup / alignment / codebook)。 |
| What is uncertain什么不确定 | ① The paper itself says the codebook loses accuracy on some task types. ② Its real-robot experiments used two cameras and it states explicitly that a third-person viewpoint makes depth perception harder — and our depth camera is head-mounted, i.e. third-person. ③ Nobody has done this on SmolVLA; the size of the change has never been measured. all untested ① 论文自己说 codebook 在某些任务类型上会掉点。② 论文的真机实验用双相机,并明确指出第三人称视角会让深度感知变难—— 而我们的深度相机正装在头上,就是第三人称。③ 没人在 SmolVLA 上做过,改动量没量过。全部未测 |
measured实测 run on this machine, raw data on disk本机跑过,有原始数据 untested未测 from a document or a paper, not verified by us照文档或论文写的,我们自己没验过 inferred推断 reasoning, not in any document推断,文档里没有
Sources:来源:
GOAL-DEPTH-INTEGRATION.md · GOAL-B0-OBJECT-TOKENS.md · GOAL-T1-DETECTOR-LATENCY.md ·
05-training/detect/owlv2/records.json · 05-training/detect/manual_review.json · 05-training/detect/bench_cameras.json
Correction, 2026-09-05 evening:2026-09-05 晚更正:
the first version of this page cited only the failed first-round detectors (OWL-ViT 27.8%, Grounding DINO 2.8%) and called detection a blocker —
while OWLv2 (zero false positives) and FT-DINOSAUR (52.4 ms) had been measured the same day and written up in GOAL-T1. That was not folded in.
The per-footage breakdown was recomputed while checking this.
本页初版只引了第一轮失败的检测器(OWL-ViT 27.8%、Grounding DINO 2.8%),把检测说成拦路的——
而 OWLv2(零假阳性)和 FT-DINOSAUR(52.4 ms)的结果同一天就测出来了,写在 GOAL-T1 里,初版没折进来。
分画面的那组数是核对这条时重新算的。