← reports index报告首页
2026-09-05 · depth integration · decision record深度整合 · 决策记录

Two ways to use depth, and how to accept object tokens深度的两条用法与物体 token 的验收

SmolVLA is the main policy, and generalisation is its biggest potential — but it is not controllable enough yet. Three additions are on the table, and they buy three different things: depth run on its own as a fallback, depth combined with the policy for accuracy, and object recognition to strengthen generalisation. This page says what each one is, what it costs, and — for the object-recognition route — how to actually accept it. SmolVLA 是主策略,泛化是它最大的潜力——但目前还不够可控。 围绕它有三件可以加的事,各买各的:深度单独跑当兜底、深度与策略结合提高准确度、物体识别增强泛化。 这一页说清每件是什么、代价是什么,以及物体识别那条到底怎么验收。

1. Three additions, three different things bought三件事,各买各的

What you add加什么 How it is used怎么用 What it buys买到什么
Depth, on its own深度,单独跑
B1 solo
No policy at all. Look once to locate the cup → IK → a scripted descend, close and lift. 完全不用策略。看一眼定位杯子 → IK → 脚本化的下降、闭合、提起。 Insurance. If SmolVLA does not work on the day, something on the table can still pick the cup up. 兜底。SmolVLA 当天不 work 的话,台上还有东西能把杯子抓起来。
Depth + SmolVLA深度 + SmolVLA
B1 handover
IK drives the arm to a pre-grasp pose over the cup; the policy does the final grasp. IK 把手臂送到杯子上方的预抓取位,策略负责最后的抓取。 Accuracy, inside the range the policy was trained on. Not generalisation — see §2. 准确度,在策略训练过的范围之内。不是泛化——见 §2。
Object recognition物体识别
B0 · object tokens
Cut the picture into objects and give the policy one token per object, alongside the existing image tokens. 把画面按物体切开,每个物体给策略一个 token,和现有的图像 token 并存。 Generalisation — a cup that looks different from the ones in the training data. 泛化——一只和训练数据里长得不一样的杯子。
Why the third one is the one that touches the potential为什么第三件才碰到那个潜力

Generalisation is what SmolVLA can do that geometry cannot — but right now the model learns an association between a patch of the picture and an action: token 137 always means the same square of the image, whatever happens to be sitting there. That is the mechanical reason a different-looking cup breaks it, and it is why generalisation currently feels uncontrollable. Object recognition changes what the model is looking at: it learns "where that object is" instead of "square 137 lit up". That is the optimisation that makes the potential controllable — not a bigger model, and not more compute. 泛化是 SmolVLA 能做而几何做不到的事——但现在模型学的是画面某一块和动作之间的关联: 第 137 号 token 永远代表画面同一个格子,不管那里坐着什么东西。 这就是「换一只杯子就失效」的机械原因,也是泛化现在显得不可控的原因。 物体识别改变的是模型看的东西:它学的变成「那个物体在哪」,而不是「第 137 号格子亮了」。 这才是让那个潜力变得可控的优化——不是更大的模型,也不是更多算力。

So: does depth need collecting, and does it go into training?所以:深度数据要不要采、进不进训练?

No, and no. Both depth uses above read the sensor live, once before the run — nothing goes into a dataset and nothing goes into the network. Object recognition does not touch depth at all; it is pure RGB, needs retraining (~1.3 h / $1) but no new recordings. 都不用。上面两种深度用法都是运行时、开跑前读一次——不入数据集,也不进网络。 物体识别根本不碰深度,是纯 RGB 的;它要重训(约 1.3 小时 / $1),但不用重采数据。

There is a route that feeds depth into training — B2, "Depth Helps" — and it is the one we are not taking, because it requires re-recording every episode in a native depth format. It is archived at the foot of this page with what it is, why it might be worth something, what it would cost and what is still uncertain. 确实有一条把深度喂进训练的路线——B2「Depth Helps」——那条我们不走, 因为它要求把每一集都用原生深度格式重录。它存档在本页页尾:它是什么、为什么可能有价值、要额外做什么、什么不确定。

2. B1 — how it works: depth outside the policyB1 用什么办法:深度在策略外

A purely geometric chain. No change to the policy, no retraining, no new data. Depth is used once before the run, not inside the control loop. 一条纯几何链路,不改策略、不重训、不要新数据。深度只在开跑前用一次,不在控制回路里。

  1. Move the head to the trial preset头移到 trial 预置位 The camera sits on a moving head, so the extrinsic is only valid at one head pose — move the head and it is void. The preset has not been saved yet. 相机装在会动的头上,外参是「头在某个姿态时」的值,头一动全废。这个预置位还没存。
  2. Detect the cup in the head RGB image开放词表检测器在 RGB 上出框 Open-vocabulary detector, prompt "a cup". Use OWLv2 at threshold 0.30 — see the card below for the numbers. 开放词表检测器,prompt "a cup"。用 OWLv2,阈值 0.30——数据见下面那张卡。
  3. Take the depth under those pixels取框内的深度像素 640×480 16-bit PNG, millimetres, outliers clipped to 100–2000 mm. 640×480 16-bit PNG,单位毫米,按 100–2000 mm 裁离群值。
  4. Back-project to a 3D point in the camera frame反投影成相机系 3D 坐标 Needs the depth camera's intrinsics. The colour camera was calibrated on 2026-09-04, but those numbers do not apply to the depth stream and must not be pasted in. What is in the code is FX = FY = 320, a guess: Z is trustworthy, X/Y approximate. 要深度相机的内参。彩色相机 2026-09-04 标定完了,但那组数不适用于深度流,不能填进去。 代码里现在是 FX = FY = 320,是猜的:Z 准,X/Y 只是近似。
  5. Transform into the arm's base frame坐标变换到臂基座 Hand-eye transform, never measured. Can be fitted offline; needs no new data. 手眼外参,没量。可以离线拟合,不需要新数据。
  6. IK to a pre-grasp pose → hand over to SmolVLAIK 摆到预抓取位 → 交棒 SmolVLA SmolVLA does the descend, close and lift. The full engineering detail of this chain — every file, every gap — is on its own page: Picking a cup with IK alone. SmolVLA 负责剩下的下降 / 闭合 / 提起。这条链路的完整工程细节——每个文件、每处缺口——单独写在了 只用 IK 抓放杯子那一页。
The detection step: what we actually measured检测这一步:我们实际测到的 measured实测

OWLv2 on our own head-camera frames, every box cropped and checked by eye: OWLv2 跑在我们自己的头相机画面上,每个框都裁出来放大人工看过:

Footage画面 Cup found找到杯子
0826 — the batch the deployed model was trained on现役模型就是用这批训的12 / 12
0903 — sep3sep39 / 12

And nothing that is not a cup gets boxed: across all 36 frames, 19 boxes, 19 of them cups — zero false positives. The VR controller and the table grommet that ruined the two earlier detectors produce no boxes at all under OWLv2. 而且不是杯子的东西不会被框出来:36 帧一共 19 个框,19 个都是杯子——零假阳性。 把之前两个检测器毁掉的 VR 手柄和桌面穿线圆孔,在 OWLv2 下一个框都不出。

Why this detection rate is good enough here. In this route the detector runs once, before the grasp starts: look, fix where the cup is, then the arm goes to that coordinate. The cup does not move during the grasp, so there is no need to keep looking. It takes 1623 ms, and waiting a second and a half before the arm moves costs nothing. 为什么这个检出率在这里够用。这条路线里检测只在开始抓之前跑一次: 看一眼,定下杯子在哪,然后手臂照这个坐标去。抓的过程中杯子不动,所以不需要一直看。 它要 1623 ms,而在手臂动之前等一秒半,不花什么代价。

Why it cannot run continuously. The robot puts out one chunk of motion every 1667 ms, and SmolVLA's own thinking already takes 1329 ms of that — leaving 337 ms. Running a 1623 ms detector every frame overruns that by nearly 5×: the arm would sit waiting on the detector and the motion would come out in stutters. Look once, then move works on this board. Keep looking does not. 为什么不能一直实时检测。机器人每 1667 ms 产出一段动作,其中 SmolVLA 自己思考就用掉 1329 ms——只剩 337 ms。 每帧都跑一个 1623 ms 的检测器,把这个预算超了将近 5 倍:手臂会一直等检测器,动作就断断续续。 看一眼再动在这块板子上成立,一直看不成立。

What we have not measured. Frames from one scene look alike, so a miss tends to be a whole-scene miss. "12 out of 12 frames" is therefore not the same as "finds the cup on every attempt". We only have three scenes — too few to say. 还没测的。同一个场景里的帧长得差不多,所以检不到往往是整场都检不到。 「12 帧里 12 帧」因此不等于「每次抓取都能找到杯子」。我们只有三个场景,太少,说不了。 untested未测

The handover collides with the 35° coverage limit交棒这个赌和 35° 覆盖直接冲突 inferred → operator's ruling, 2026-09-05推断 → 操作者 2026-09-05 判定

B1's handover version exists to buy new cup positions. But moving the arm to a new position puts shoulder_pan outside the trained range — the right arm's demonstrations cover only 35.1° of lateral pan (std 5.15°), measured at the frame where the gripper closes. The state SmolVLA receives at handover therefore lands in exactly the dimension it has never seen. B1 的交棒版存在的目的是买换位置。可把手臂开到新位置,就意味着 shoulder_pan 被推到训练范围之外—— 右臂示范在夹爪闭合那一帧测出的横向覆盖只有 35.1°(std 5.15°)。 SmolVLA 交棒时收到的状态,正好落在它没见过的那一维上。

Operator's ruling, 2026-09-05: this combination is not the answer to new positions. After the handover it is still SmolVLA driving; a position it has never seen stays out of reach however precisely the arm is placed. 操作者 2026-09-05 判定:这个组合不是新位置的答案。交棒之后跑的还是 SmolVLA, 它没见过的位置,手臂摆得再准也够不着。

The five-minute handover test is still worth doing at a cup position inside the trained band — it tests whether the handover mechanism works at all, which matters for the reliability question. What it must not be sold as is a route to position generalisation. That question is treated in full on the IK page, §3. 那个五分钟的交棒测试仍然值得做,但要在训练带之内的杯位上做——它验的是交棒机制本身能不能接上, 这对「更有把握」那个问题有意义。不该被当成通往位置泛化的路子。 这个问题在 IK 那一页 §3 有完整讨论。

What B1 is actually for: a purely geometric grasp, as demo insuranceB1 现在真正的用途:纯深度抓取,当兜底 demo not written · not run未写 · 未跑过

The same geometric chain, but carried all the way through — descend, close and lift all driven by IK and a script, never handing over to the policy. Position generalisation is then free: geometry needs no training coverage, and the 35° limit simply does not apply. 同一套几何链路,但一路做到底——下降、闭合、提起全部走 IK + 脚本,中途不交给策略。 这样位置泛化就是 免费的:几何不需要训练数据覆盖,35° 那条限制根本不存在。

What it buys: insurance for the demo. If the policy does not work on the day, there is still something on the table that can pick the cup up. This is a fallback, not a policy improvement — it does not make the model one byte better. 它买的是什么:demo 的保险。策略当天不起效果时,台上还有东西能把杯子抓起来。 这是兜底,不是策略改进——它一个字节也没改善模型。

It costs more than the handover version, not less. With no policy at the end to absorb error, accuracy rests entirely on perception and calibration. None of the geometry prerequisites can be skipped, and the tolerance is tighter. 代价比交棒版更高,不是更低。末端没有策略帮忙兜误差,精度全压在感知和标定上。 几何那几项前置一项都不能省,而且对精度的要求更严。

Detection is no longer the blocker. OWLv2 has zero false positives and finds the cup in every frame of the current footage, and 1623 ms is fine because this route also looks once before moving. The real unknowns are calibration accuracy and the grasp sequence, which has not been written. Full breakdown here. 检测这一关不再是拦路的。OWLv2 零假阳性,在当前画面上每一帧都找到了杯子; 1623 ms 也不碍事,因为这条路同样是动之前看一眼。真正的未知在标定精度,以及那段还没写的抓取序列。 完整拆解在这一页。

3. B0 — how it works: object tokensB0 用什么办法:物体 token

Instead of computing a 3D coordinate, cut the picture into objects, give each one a token appended to the prefix, and let the model decide which to look at. No depth, no intrinsics, no hand-eye, no new recordings — but it does need retraining. 不算 3D 坐标,把画面按物体切开,每个物体给一个 token 接到 prefix 末尾,让模型自己挑该看谁。 不用深度、不用内参、不用手眼、不用重采数据,但要重训。

  1. Use FT-DINOSAUR for the object slots — settled 2026-09-05用 FT-DINOSAUR 出物体 slot——2026-09-05 已经定了 small@224 runs in 52.4 ms per frame (median, N=50; p95 53.8, 152 MB of VRAM) and outputs 7 fixed slots × 256 dims — so the fixed-count constraint below is satisfied for free. Quality has passed a by-eye review: median coverage 94%. 518 resolution is not better than 224 and can be dropped. This overturned the earlier guess that "it is a ViT, it cannot be much faster than 177.7 ms" — it measured 3.4× faster than that supposed floor.
    OWLv2 is ruled out for this job: 1623 ms overruns the 337 ms per-cycle budget by 4.8×, and its per-frame detection rate would leave many frames with no object token at all, which would teach the model to ignore the input. FT-DINOSAUR always emits 7 slots, so that problem does not arise. The division of labour is pay once for the expensive one, run the cheap one every frame: OWLv2 at the start to identify which slot is the cup, FT-DINOSAUR to track it. Memory was measured too: detector + SmolVLA + geometry all resident is 3190 MB RSS, with 1844 MB still free. measured + by-eye review
    small@224 单帧 52.4 ms(中位,N=50;p95 53.8,显存 152 MB),输出 7 个固定 slot × 256 维—— 下面第 5 条那个「数量必须固定」的约束自动满足。质量已过人工看图:覆盖率中位 94%。 518 分辨率不比 224 好,可以排除。这推翻了此前的推算(「它是个 ViT,不可能比 177.7 ms 快多少」)——实测比那个所谓下界还快 3.4 倍。
    OWLv2 在这个位置被否掉:1623 ms 超每周期 337 ms 预算 4.8 倍;而且按它的每帧检出率, 会有很多帧一个物体 token 都没有,那等于教模型忽略这个输入。FT-DINOSAUR 恒出 7 个 slot,没有这个问题。 分工是贵的付一次、便宜的每帧跑:开局用 OWLv2 认出哪个 slot 是杯子,之后用 FT-DINOSAUR 跟踪它。 内存也测了:检测器 + SmolVLA + 几何三者同时驻留 RSS 3190 MB,系统仍剩 1844 MB。实测 + 人工看图
  2. Add an object_proj加一个 object_proj Copy state_proj (nn.Linear(32, 960), modeling_smolvla.py:508) to project the object features to 960 dims. 照抄 state_proj(nn.Linear(32, 960),modeling_smolvla.py:508),把物体特征投到 960 维。
  3. Append [B, N_obj, 960] at the end of embed_prefix在 embed_prefix 末尾 append [B, N_obj, 960] modeling_smolvla.py:552 is just a list that gets concatenated at the end. About 15 lines of change. modeling_smolvla.py:552 那里就是一个 list 最后 concat。改动约 15 行。
  4. att_mask must be 1, and the tokens must go lastatt_mask 必须标 1,且必须放在最后 This is the only safe way to write it. att_masks is not an on/off switch for visibility — it is a block boundary marker: a cumulative sum gives the block number, and token i can see token j iff block(j) ≤ block(i). Marking 1 at the end opens a new block, which leaves the image and language outputs bit-for-bit unchanged. Inserting before the state token with a 0 instead rewrites the pretrained visual and language features. verified in the source 这是唯一安全的写法。att_masks 不是「能不能看」的开关,是块边界标记: cumsum 给出块号,token i 看得到 j ⟺ 块号(j) ≤ 块号(i)。标 1 放最后 = 开新块 = 图像和语言的输出逐比特不变。 插在 state 之前且标 0,会改写预训练的视觉 / 语言特征。读过源码
  5. N_obj must be fixedN_obj 必须固定 att_masks is a python list shared across the whole batch. When fewer objects are found, pad with pad_masks=0 — padding tokens are masked on both the key and the query side, so empty objects contaminate nothing. att_masks 是 python list,整个 batch 共享同一长度。检出不足时用 pad_masks=0 补—— padding token 在 key 和 query 两侧都被屏蔽,空物体不污染任何东西。
  6. Train only the action expert and object_proj只训 action expert + object_proj The deployed config is already freeze_vision_encoder: True and train_expert_only: True — SigLIP does not move a parameter, so this is compatible with the existing training mode. Object tokens are computed offline from the RGB we already recorded, so no re-recording. About 1.3 h / $1. 部署配置本来就是 freeze_vision_encoder: True、train_expert_only: True,SigLIP 一个参数不动, 和现有训练模式兼容。物体 token 从已录的 RGB 离线算,所以不用重采数据。约 1.3 小时 / $1。
Two things to settle up front两件事先说死

Speed is not the reason to do this. SmolVLA takes 1329 ms per inference and the three camera streams are only 13% of it (177.7 ms). Cutting tokens saves a fraction of that, and the segmentation model's own cost eats into it — the net could well be negative. The only argument that stands up is generalisation from object-level structure. measured 速度不是理由。SmolVLA 一次推理 1329 ms,三路视觉只占 13%(177.7 ms)。 token 砍下来省的那点,再减去分割模型自己的开销,净收益很可能为负。 唯一站得住的理由是物体级结构先验带来的泛化。实测

Zero-initialisation is not identity. Even with object_proj's last layer zeroed and every value zero, the extra keys still take part in the softmax normalisation and dilute the existing attention weights. A warmup or a small initial learning rate is still needed. 零初始化不等于恒等。就算 object_proj 最后一层置零、value 全是零, 多出来的 key 仍然参与 softmax 归一化、会稀释原有的注意力权重。仍需 warmup 或小学习率起步。

4. Acceptance and A/B for B0 — this section did not existB0 的验收与 A/B——这一节原本没有

The B0 document lists work items T1/T2/T3, and T3 says "compute the net benefit and decide" — with no criterion and no controlled design. Here it is. Three hard constraints shape it: B0 那份文档列了 T1 / T2 / T3,T3 写的是「算净收益并决定」——没有判据,也没有对照设计。 下面补上。先说三条决定设计形状的硬约束:

  1. Only appearance generalisation can be tested只能测「换外观」 B0 addresses a different-looking cup in the same place. It cannot buy new positions — those are bounded by the 35° coverage and only more data fixes that. Using overall success rate as the metric tests the wrong thing. B0 对口的是同一个位置换一只长得不一样的杯子。它买不到换位置——那受 35° 覆盖限制,只能补数据。 拿总成功率做 A/B 是在测错东西。
  2. Real-robot success rates need far more trials than we can run实机成功率的样本量根本不够 Resolving a 20-percentage-point difference needs about 90 trials per arm (Fisher, two-sided α = .05, power .8). We already recomputed Oat-VLA's 29/49 vs 20/49 as p = 0.106, not significant — running 20 trials would put us in the same hole. 要在成功率上看出 20 个百分点的差,每组约需 90 次(Fisher 双侧 α=.05,power .8)。 我们自己算过 Oat-VLA 的 29/49 vs 20/49 是 p = 0.106 不显著——跑 20 次会掉进同一个坑。
  3. The baseline has to be retrained基线必须重训 Arm A cannot be the deployed checkpoint, or the difference also contains config, step count and seed. A = same config, same seed, same steps, the only difference being no object tokens. A 组不能拿现役 checkpoint 顶替,否则差异里混着训练配置、步数、seed。 A 组 = 同 config、同 seed、同步数,唯一区别是不加物体 token。

Then three gates, cheapest first. If one fails, stop — do not go on to the next. 然后是三道门,便宜的在前。前一道不过就停,不进下一道。

GATE 1门 1Negative control负对照free · offline零成本 · 纯离线

After training B0, replace every object token with zeros at inference (then repeat with random vectors). 训完 B0 之后,推理时把物体 token 全部换成零向量(再跑一遍随机向量)。

Why this goes first:为什么先做它: it catches the most likely failure — the interface is correct, training completed, and the model quietly learned to ignore the new tokens. On the robot that looks identical to "same as baseline" and would be misread as "B0 does not help". 它专抓最可能的失败模式——接口写对了、训练也跑完了,但模型学会了忽略新 token。 这种情况在实机上表现为「和基线一样」,会被误读成「B0 没用」。
GATE 2门 2Offline, new cup appearance离线换外观~10 episodes · test set only采 ~10 集 · 只做测试集
Criterion, fixed in advance:事先写死的判据: B's error on the new cups must be lower than A's, and that drop must exceed the difference between the two on the old cups — the gain has to come from the appearance change, not from fitting better overall. Fail → stop, and write "do not do it" into T3, which the document explicitly allows. B 在新杯子上的误差比 A 低,且这个降幅大于两者在旧杯子上的误差差—— 收益必须来自换外观,不是整体拟合更好。不过 → 停,把「不做」写进 T3,文档明确允许这个结论。
GATE 3门 3Paired trials on the robot实机配对试验only if gate 2 passed门 2 过了才做
Fixed in advance:事先写死: primary metric is the paired difference at the "lifted" level; secondary is the ordinal score. n and the stopping rule go into the document before the first run — no adding trials after looking at the data. 主指标 = 「提起」这一级的配对差;次指标 = 序数分。 n 和停止规则开跑前就写进文档,不许看着数据加试验。
One cheap thing worth checking first顺带一条几分钟就能查完的

An open question from the B0 document: how many different cups appear across the existing 202 episodes? If only one, the model has probably locked onto that cup's appearance and B0 has a lot to gain; if several, the model may already have some appearance invariance and there is less to buy. Purely offline — read the keyframes of the merged dataset. B0 文档末尾留的一个问题:现有 202 集里到底用了几种杯子? 全程只有一只 → 模型很可能已经锁死在那只杯子的外观上,B0 的收益空间大; 本来就有好几只 → 模型可能已经有外观不变性,能买到的就少。纯离线——读合并集的关键帧就行。

5. B2 · Depth Helps — archived, not doing itB2 · Depth Helps 式注入(不做 · 存档)

Detail内容
What it is它是什么 Freeze SmolVLA's SigLIP backbone and add two small modules alongside it: a Depth Completion Module that predicts depth features from RGB via cross-attention, and a Depth-Aware Codebook that discretises depth for noise robustness. Only those two modules are trained. Source: arXiv 2408.05107 把 SmolVLA 的 SigLIP 主干冻住不动,在旁边加两个小模块: Depth Completion Module(用 cross-attention 从 RGB 预测深度特征)和 Depth-Aware Codebook(把深度离散化,抗噪)。 只训这两个小模块。出处 arXiv 2408.05107
Why it might
be worth it
为什么可能
有价值
It is the only one of the three that actually feeds depth into the policy. The paper reports 57.95% → 63.95%, with the largest gain on long-horizon tasks (~12%); it needs only about 20 real trajectories per task and we have 30. And at inference it can run on RGB alone (63.15% vs 63.95%, a 0.8-point drop) — depth is only needed during training, so deployment gets simpler, not harder. All three points fit our situation closely. 它是三条里唯一真正把深度喂进策略的。论文报 57.95% → 63.95%,长时程任务提升最大(约 12%); 每个任务只要约 20 条真机轨迹,而我们有 30 集。而且推理时可以只用 RGB(63.15% vs 63.95%,只掉 0.8%)—— 深度只在训练时需要,部署反而更省事。这三点和我们的处境高度吻合。
What extra it
would take
需要额外
做什么
Re-recording. Depth has to go into the dataset as a native feature (orbbec_depth_camera.py already supports it), and the 30 episodes recorded on 9-02 / 9-03 are still in the old sidecar format and cannot be used. This is the only large cost among the three routes, and it is exactly why the route was shelved. It also needs three-stage training (warmup / alignment / codebook). 重录数据。深度必须以原生特征录进数据集(orbbec_depth_camera.py 已打通), 而 9-02 / 9-03 那 30 集还是旧的旁路格式,用不了。这是三条路线里唯一一笔大成本,也正是它被搁置的原因。 另外要走三阶段训练(warmup / alignment / codebook)。
What is uncertain什么不确定 ① The paper itself says the codebook loses accuracy on some task types. ② Its real-robot experiments used two cameras and it states explicitly that a third-person viewpoint makes depth perception harder — and our depth camera is head-mounted, i.e. third-person. ③ Nobody has done this on SmolVLA; the size of the change has never been measured. all untested ① 论文自己说 codebook 在某些任务类型上会掉点。② 论文的真机实验用双相机,并明确指出第三人称视角会让深度感知变难—— 而我们的深度相机正装在头上,就是第三人称。③ 没人在 SmolVLA 上做过,改动量没量过。全部未测

6. Related page相关页面

Picking a cup with IK alone只用 IK 抓放杯子 The engineering detail behind §2: the full chain from camera pixel to closed gripper with no policy in it, each step naming the file that implements it or saying it does not exist. Includes what is calibrated and what is not, and a section on whether handing over to SmolVLA for the final grasp makes it more reliable inside the trained range. §2 背后的工程细节:从相机像素到夹爪闭合的完整链路,中间不放策略,每一步点名实现它的文件或直说它不存在。 包括哪些标定完了哪些没有,以及一节专门讨论:最后的抓取交给 SmolVLA,在训练范围内能不能更有把握。

7. Evidence levels证据等级

measured实测 run on this machine, raw data on disk本机跑过,有原始数据 untested未测 from a document or a paper, not verified by us照文档或论文写的,我们自己没验过 inferred推断 reasoning, not in any document推断,文档里没有

Sources:来源: GOAL-DEPTH-INTEGRATION.md · GOAL-B0-OBJECT-TOKENS.md · GOAL-T1-DETECTOR-LATENCY.md · 05-training/detect/owlv2/records.json · 05-training/detect/manual_review.json · 05-training/detect/bench_cameras.json

Correction, 2026-09-05 evening:2026-09-05 晚更正: the first version of this page cited only the failed first-round detectors (OWL-ViT 27.8%, Grounding DINO 2.8%) and called detection a blocker — while OWLv2 (zero false positives) and FT-DINOSAUR (52.4 ms) had been measured the same day and written up in GOAL-T1. That was not folded in. The per-footage breakdown was recomputed while checking this. 本页初版只引了第一轮失败的检测器(OWL-ViT 27.8%、Grounding DINO 2.8%),把检测说成拦路的—— 而 OWLv2(零假阳性)和 FT-DINOSAUR(52.4 ms)的结果同一天就测出来了,写在 GOAL-T1 里,初版没折进来。 分画面的那组数是核对这条时重新算的。