2026-09-04 · written for anyone joining or returning to the project
Every number here is either measured and reproducible from this repository, or explicitly marked as not measured. Incorporates the 2026-09-04 remote-task results, including one correction that invalidated a figure repeated across several earlier documents — see §7.4. Where a figure comes from someone else's paper or product page, it says so.
A bartender robot whose selling point is not throughput. Commercial bartender robots already make drinks quickly and consistently — Makr Shakr Veloce does 240 drinks/hour — and they do it by building the entire installation around the machine: bottles on a fixed rail, glasses in a fixed jig, arms replaying pre-programmed trajectories. Their AI sits in the recommendation layer. The arm does not look at anything.
We are building the one that watches the customer: greets them, makes the drink, rotates the glass so they can see it, and hands it toward where they actually are. That requires learned manipulation rather than recorded trajectories, and it requires perception directed at a person, not just at an object.
The honest state of that ambition: our grasp success rate is currently 8/83 (95 % CI 4.3–18.1 %), and we have not recorded a single pouring demonstration.
| Compute | Jetson Orin Nano Super, 7.4 GiB shared CPU/GPU memory, 16 GB swap |
|---|---|
| Arms | 2 × 6 DoF, Feetech serial-bus servos |
| Head | Pan/tilt (servos 7, 8 on the left arm bus) |
| Cameras | 3 × RGB — head 424×240, left wrist 640×480, right wrist 640×480 |
| Depth | Orbbec, mounted on the head |
| Gripper | Self-designed compliant fin-ray, 3D printed |
| Base | 3 velocity dimensions, currently unused (always zero) |
| Not present | Force / torque sensing. Servo load is not in the observation. |
Action and state are 17-dimensional: 6 + 6 arm, 2 head, 3 base velocity.
Why the compute matters strategically: 7.4 GiB cannot hold a 7-billion-parameter VLA. That rules out the OpenVLA family and forces us toward smaller models and cleverer representations rather than bigger ones. It is a constraint, but it is also what makes any result we get transferable to affordable hardware.
| Total episodes on the box | 298 (116,429 frames, 30 fps) |
|---|---|
| Of which marked trash/soak | 42 |
| Usable | 256 |
Under the team org (xlerobot-team) | 252 |
Largest sets: xlerobot-cup-grasp-20260820-0230 (50 ep — what the current policy was trained on), xlerobot-right-pick-cup-20260826-0042 (40 ep — current gripper), xlerobot-left-pick-cup-sep3-20260903-2015 (29 ep — most recent).
All of it is cup grasping. There are zero demonstrations of pouring, shaking, stirring, presenting or handing over.
Depth exists as a side-channel: <dataset>/depth/ep<N>-<timestamp>/frame_NNNNNN.png, 640×480 16-bit PNG in millimetres, ~863 MB total.
It is not in any dataset's feature list, so no training or inference code can read it. Measured quality where it does exist: 76–86 % valid pixels, median depth ≈ 500 mm — the data itself is good. First 1–2 frames per episode are zeros (sensor warm-up); there are outliers up to 32 m that are noise and must be clipped.
Native in-dataset depth recording is implemented and waiting on a service restart.
head_1 σ = 0.0000 range [-4.308, -4.308]
head_2 σ = 0.0000 range [49.846, 49.846]
Across all 19,453 frames of the training set, the head never moved. The policy has never seen a moving camera. This is why gaze must stay frozen during manipulation and can only follow a person before and after — see §7.
Seven checkpoints exist locally. Trial results across 37 result files / 114 trials:
| Checkpoint | Measured | Batches |
|---|---|---|
smolvla_20260820-1502/016000 (= last) | 6/50 (12.0 %, 95 % CI 4.5–24.3 %) | 25 |
a100-6000/006000 | 0/23 (0 %, 95 % CI 0–14.8 %) | 3 |
…cup-grasp-right-b8-u20k/020000 | 0/4 (0 %, 95 % CI 0–60.2 %) | 2 |
…right-pick-b16-3cam-u20k/020000 | never tested | 0 |
smolvla_20260820-1502/018000 | never tested | 0 |
act_20260820-0429/020000 | never tested | 0 |
act_20260820-0429/004000_from_hf | never tested | 0 |
Overall: 8/83 counted trials succeeded — 9.6 %, 95 % CI 4.3–18.1 %.
Two things follow that are easy to get wrong:
0/4 is not a failed model, it is an untested one. Its 95 % upper bound is 60.2 %.b16-3cam-u20k, which was trained on all three cameras and is therefore not comparable to the others in input form.Zero of the 37 batches were pre-registered (a success criterion fixed before the run). By our own standard, all existing trial data is exploratory.
| Status | |
|---|---|
| Bimanual teleoperation and recording (VR and leader-arm) | ✅ |
| Leader-follower interactive recording (port 8781) | ✅ |
| Joint-level safety layer | ✅ rejected-frame rate tuned 96.6 % → 0.49 % |
| Control panel: jobs, datasets, checkpoints, trials, calibration (port 8770) | ✅ |
| Face detection + expression + OLED expressive face + dialogue | ✅ as a subsystem |
Live head-camera viewer with face + cup boxes, in a browser (cup_test.py, port 8795) | ✅ contributed by Ville — see §5.1 |
Head tracking of a primary face (face_follow.py) | ✅ not connected to the arm |
| Depth capture | ✅ side-channel only, see §3.1 |
| Trial records: schema v2, Clopper-Pearson intervals, failure modes, funnel, pre-registration | ✅ new 2026-09-03/04 |
| File drop onto the robot (port 8790) | ✅ new 2026-09-04 |
03-software/conversational_ai/cup_test.py (165 lines) serves the head camera in a browser with face and cup boxes drawn on it: ./run.sh cup_test.py, then open port 8795. It uses YuNet for faces — the same choice we reached independently — and EfficientDet-Lite0-int8 for cups.
Two design details worth keeping: it finds the Orbbec camera by name rather than by a /dev/videoN number, and it falls back to the shared frame in /dev/shm when another program holds the camera. Both directly avoid failure modes this project has hit before.
One caveat that matters (§6.1): it opens the camera at 1280×720, whereas the head camera is recorded — and therefore seen by the policy — at 424×240. Nine times the pixels. A detector that looks convincing in the live viewer may still fail on the frames the policy receives, which is exactly what our offline evaluation found for its default cup model.
A useful corollary: the camera evidently can deliver 1280×720, so raising the recording resolution may be a configuration change rather than new hardware. <span>UNTESTED — bandwidth, disk and retraining cost all unknown</span>
| Reliably locating the cup | 🔧 Was four failures; OWLv2 changed that on 2026-09-04. A zero-error configuration now exists on 36 reviewed frames, but scene generalisation is untested and the acceptance criterion needs rewriting. Earlier history: Three hand-written detectors (79 / 97.5 / 97.6 % detection, jitter to 0.9 mm) were all locking onto the robot's own arm. On 2026-09-04 four pretrained open-vocabulary detectors were evaluated offline with human review of every frame — best result 5/36 = 13.9 % on the target cup, against an acceptance bar of 90 %. See §6.1 |
|---|---|
| Camera intrinsics | ❌ FX=FY=320 is a guess |
| Hand-eye extrinsics | ❌ never done |
| Depth↔RGB time alignment | ❌ never verified |
trial head preset | ❌ does not exist, and everything geometric depends on it |
| Any policy reading depth | ❌ SmolVLA, ACT and every lerobot policy are RGB-only |
| Face subsystem connected to the arm | ❌ |
| Identity recognition | ❌ — confirmed 2026-09-04: the face stack is YuNet, a detector. No identity anywhere |
| Pouring / shaking / stirring / presenting | ❌ no demonstrations, no policy |
| Keyframe capture in trials | ✅ implemented 2026-09-04, --dry verified |
Four pretrained detectors, run offline over recorded frames, every frame reviewed by eye:
Revised twice on 2026-09-04. First downward, after every box was cropped and identified by eye (contact-sheet thumbnails had made a VR controller look like a cup). Then substantially upward, after switching to OWLv2.
| Detector | Boxes | Actually a cup | The target cup | Latency |
|---|---|---|---|---|
| EfficientDet-Lite0 int8 | 0/36 | — | — | 66 ms CPU |
| EfficientDet-Lite2 | 7/36 | not itemised | — | 411 ms CPU |
| OWL-ViT base-patch32 | 16/36 | 10 (62.5 %) | 5 | 225 ms GPU |
| Grounding DINO tiny | 27/36 | 1 | 1 | 2657 ms GPU |
OWL-ViT's false positives were a VR controller left on the table (5 boxes); Grounding DINO's were cable-grommet holes in the table, the mouse, the gripper and cabling (26 of 27).
Same 36 frames, same queries, same selection rule; only the model changed.
| Configuration | Boxes | Target cup | Other cup | Not a cup | Verdict |
|---|---|---|---|---|---|
| OWL-ViT @0.05 | 16 | 5 | 5 | 6 | 62.5 % are cups |
| OWLv2 @0.05 | 24 | 18 | 3 | 3 | 87.5 % are cups |
| OWLv2 @0.30 | 19 | 16 | 3 | 0 | 100 % are cups |
| OWLv2 @0.30 + background difference @50 % | 15 | 15 | 0 | 0 | first zero-error configuration |
The VR controller and the table holes produce zero boxes under OWLv2.
And thresholding went from useless to decisive, which is the structural change:
| Target-cup scores | Non-cup scores | Separable by threshold? | |
|---|---|---|---|
| OWL-ViT | 0.051 – 0.125 | 0.058 – 0.213 | ❌ the false positives score higher |
| OWLv2 | 0.18 – 0.71 | 0.09 – 0.25 | ✅ 0.30 cuts cleanly |
| Latency | 1623 ms/frame — 7× OWL-ViT, and slower than a full SmolVLA inference (1329 ms). Per-frame online use is impossible. Viable as a staged call: "stop, look once, lock the target" |
|---|---|
| Per-frame detection rate | 41.7 % (15 boxes from 36 frames) |
| Scene generalisation | ❌ not tested at all. 36 frames, one room, one lighting condition |
| The background reference | A cross-episode pixel median, not a real empty-table photograph — and it is contaminated by motion blur from the arm. So the layer's measured performance is a lower bound |
Background subtraction requires the reference image and the runtime camera pose to match. The head is locked during trials, so within a session this holds. But changing the head preset invalidates the reference, and nothing will warn you. This must go into the operating procedure.
| Intervention | False positives | Target cups | Net | Previously called | |---|---|---|---|---| | OWL-ViT, drop "a drinking glass" | 6 → 7 (worse) | 5 → 1 | loss | correctly, a loss | | Grounding DINO, drop "a drinking glass" | 26 → 1 | 1 → 1 (no loss) | clear gain | wrongly, "over-filtered" | | OWL-ViT → OWLv2 | 6 → 3 | 5 → 18 | double win | — | | OWLv2 threshold 0.05 → 0.30 | 3 → 0 | 18 → 16 | worthwhile | — |
The Grounding DINO run had been dismissed because only the box count was reported — it dropped to 2/36 and that looked like a collapse. Counting both numbers shows it removed 25 false positives at zero cost. Reporting one number instead of two produced a wrong conclusion twice in one investigation.
| Layer | Was meant to remove | Left to remove after OWLv2 | Its value now | |---|---|---|---| | ① Semantic (open-vocabulary + corrected distractors) | arm, mouse, cables | — | non-cup false positives are zero | | ② Depth plane residual | flat things (the table holes) | 0 — the holes no longer produce boxes | ≈0 as a filter, but still the only source of height/distance for grasping | | ③ Empty-table background difference | permanent scene objects | 3 (the other plastic cup) | the only layer still removing errors |
New order: ① → ③ → ②. This is not a rejection of depth — its role changed from filtering to supplying 3D position, and that still needs intrinsics and depth/RGB alignment, which are on-site work. It should no longer block validating layer ③.
| # | Stage | Head | Arm |
|---|---|---|---|
| 1 | Greet, converse | follows the face | still |
| 2 | Make the drink | locked | policy |
| 3 | Rotate the glass to show it | follows the face | policy |
| 4 | Push it toward the customer | follows the face | policy |
Two independent constraints force the middle row: the policy has never seen a moving head (§3.2), and the depth camera's extrinsics are only valid for one head pose. This was not chosen; it is what survived elimination.
Stage 1 is the cheapest demonstrable win in the project — the arm is still, so it conflicts with nothing, and no commercial product does it.
Handing a glass into someone's hand is a human-robot handover problem whose core difficulty is when to release. Too early and it drops; too late and the person cannot take it. Humans use touch. We have no force sensing (§2). Pushing the glass along the table keeps support under it the whole time and needs no release decision.
This is the project's one real technical bet, and it is unproven — see §4.
A worked example from 2026-09-04. Several documents in this repo — including an earlier draft of this one — stated "SmolVLA runs at 146 ms/frame (6.9 Hz)". Re-measuring produced 1329 ms, nine times larger. Tracing it back: the 146 ms figure was a real measurement from 2026-08-21, but of ACT, not SmolVLA; the model name was substituted somewhere in the chain of hand-off documents. Re-measuring both on the same machine confirmed it: ACT 129.5 ms, SmolVLA 1329.2 ms.
A second, framework-level error travelled with it: "6.9 Hz cannot keep up with 30 Hz training" is not a meaningful statement about a chunked policy. SmolVLA plans 50 action steps per inference, which is 1667 ms of motion at 30 Hz. The right question is whether one inference fits in 1667 ms — it does, at 1329 ms, with 337 ms to spare.
Both errors survived multiple documents because nobody re-measured. That is the failure mode this rule exists to catch.
measured / not measured / tried but blocked. Adopted after a review found untested inferences being reported as facts. It applies to other people's papers too: Oat-VLA reports 59 % vs 41 % on real robots; recomputing from its own counts gives 29/49 vs 20/49, overlapping intervals, Fisher p = 0.106 — not significant. We cite its token reduction, not its success rate.
| Conflict | Status | |
|---|---|---|
| 1 | face_follow.py and the policy both own the left serial bus, and the safety layer overwrites the head every frame. They cannot run simultaneously | Resolution chosen: head control stays inside the policy process, only the target source changes per phase. Not implemented |
| 2 | The safety layer has no concept of a person. Human safety today = a fixed geometric envelope + someone holding the e-stop. "Reach toward the customer" is the exact motion that envelope is shaped to prevent | Needs a parallel design. Not started |
| 3 | The per-frame step cap was derived from grasping data (wrist_roll ≤ 10.4°/frame) and it silently slows motion rather than rejecting it. A batch of presenting demonstrations could be recorded flattened, invisibly | Mitigation: measure the motion once before recording. Pending on-site |
| 4 | Three face detectors would exist if Ville's is added (Haar in scripts/, YuNet in conversational_ai/, his). Convergence needed | Open |
One factual correction found on 2026-09-04: policy_safety.py's own note asserts the training set carries head_motor_2 = 99.649, and concludes from that arithmetic that the head sat 564 counts outside the calibrated range — which is why the absolute clamp skips the head entirely. The data says 49.846, which recomputes to inside the range by 2 counts. The premise is wrong. Harmless while the head is pinned; not harmless once the head moves.
Verified from vendor pages and named sources (full analysis with images: 04-reports/market-analysis/, not published publicly):
| System | Form | Throughput | Price |
|---|---|---|---|
| Makr Shakr Toni | 2 arms, 158 bottles | 80+/hr (official) | — |
| Makr Shakr Veloce | 2 arms, speed-oriented | 240/hr | — |
| Cecilia.ai | Screen persona + pumps, voice dialogue | 120/hr | $45,000 or $2,000/mo |
| Yanu (Tartu, Estonia) | Arm inside an enclosed pod, dialogue + payments | 100/hr | — |
| MIXO / TendedBar / Botrista | Pumps, carousels | — | industry band $30k–$100k+ |
The gap, stated narrowly enough to survive checking: conversational bartender robots exist (Cecilia, Yanu) and expressive robot arms exist (Makr Shakr), but no product puts both in the same body — and in all of them the arm motion is pre-programmed. Their sourced description: "The recipe database contains pre-programmed ratios, sequences, and pour volumes", while machine learning "tracks popular combinations, personalizes suggestions". The AI is in the recommendation layer, not the motion layer.
| Role | Location | Reaches the robot how |
|---|---|---|
| Project lead | China | Tailscale (100.107.145.111) |
| Face recognition & interaction (Ville) | Taiwan (exchange) | Not yet on the network — an invite would let him measure on the real board |
| On-site data collection & hardware | With the robot | Physically |
Only one person can touch the machine. Everything else is remote, which is why the on-site brief (BRIEF-ONSITE-HARDWARE.md) is written around "anything you don't measure while you are there, nobody can measure later".
Since detection failed on disambiguation rather than recognition, four routes exist. Recommendation from the analysis: (a).
| Route | Assessment | |
|---|---|---|
| a | Use depth to disambiguate | Recommended — depth is the only signal that is insensitive to appearance, and we already have 863 MB of it. Blocked on camera intrinsics and depth/RGB alignment, which are on-site work |
| b | Raise head-camera resolution | 424×240 is genuinely small, but this does not address ranking |
| c | Better language prompts | Cheap to try; unlikely to fix ranking |
| d | Fiducial markers on the cups | Reliable fallback, but concedes the generalisation claim |
trial head preset — 1 minute, everything geometric depends on itface_follow.py was found to be non-functional — it was never "two working options". Head-clamp limits are now derived (head_motor_1 ±149.98°, head_motor_2 ±50.02°) but deliberately not applied yet; recommendation is to land them together with head-following rather than in isolation--dry verified, trials_io reads the new fields end to endNothing downstream matters until the cup can be reliably located. Pouring, presenting and handing over all begin with a successful grasp, and grasping currently succeeds 8 times in 83.
| Risk | First mitigation | |
|---|---|---|
| 1 | Grasping never gets good enough. 8/83 today. Detection failed four times, then OWLv2 produced the first zero-error configuration (15/15 on reviewed frames). The risk eased but did not close: one room, one lighting, 36 frames, and per-episode hit rate never measured | Widen to ~150 frames across batches to rule out luck; then re-measure per episode |
| 2 | Demonstrations cannot teach "elegant". Pour rate, stop timing and style exist only in the demonstrations, and we have none. Better perception does not tell a model how much to pour | Measure the motion first (§8 item 3), then record deliberately |
| 3 | Compute runs out — now quantified. One SmolVLA inference is 1329 ms and plans 50 steps = 1667 ms of motion at 30 Hz — 80 % of budget used, 337 ms margin. Any per-frame detector eats that margin | Measured 2026-09-04. Grounding DINO tiny alone is 2657 ms, so online use is out |
All three have a first step that is remote, cheap, and can be done before the next batch of demonstrations is recorded.
Companion documents: BRIEF-FACE-INTERACTION.md · BRIEF-ONSITE-HARDWARE.md · HANDOFF-REMOTE-2026-09-04.md · TOMORROW-2026-09-04.md · three GOAL files · 04-reports/market-analysis/ · STATUS.md (machine-generated, always current).