← reports index
2026-09-04 · complete status · incl. OWLv2 results

XLeRobot Bartender — Complete project status

2026-09-04 · written for anyone joining or returning to the project

Every number here is either measured and reproducible from this repository, or explicitly marked as not measured. Incorporates the 2026-09-04 remote-task results, including one correction that invalidated a figure repeated across several earlier documents — see §7.4. Where a figure comes from someone else's paper or product page, it says so.


1. What we are building, in one paragraph

A bartender robot whose selling point is not throughput. Commercial bartender robots already make drinks quickly and consistently — Makr Shakr Veloce does 240 drinks/hour — and they do it by building the entire installation around the machine: bottles on a fixed rail, glasses in a fixed jig, arms replaying pre-programmed trajectories. Their AI sits in the recommendation layer. The arm does not look at anything.

We are building the one that watches the customer: greets them, makes the drink, rotates the glass so they can see it, and hands it toward where they actually are. That requires learned manipulation rather than recorded trajectories, and it requires perception directed at a person, not just at an object.

The honest state of that ambition: our grasp success rate is currently 8/83 (95 % CI 4.3–18.1 %), and we have not recorded a single pouring demonstration.


2. Hardware

ComputeJetson Orin Nano Super, 7.4 GiB shared CPU/GPU memory, 16 GB swap
Arms2 × 6 DoF, Feetech serial-bus servos
HeadPan/tilt (servos 7, 8 on the left arm bus)
Cameras3 × RGB — head 424×240, left wrist 640×480, right wrist 640×480
DepthOrbbec, mounted on the head
GripperSelf-designed compliant fin-ray, 3D printed
Base3 velocity dimensions, currently unused (always zero)
Not presentForce / torque sensing. Servo load is not in the observation.

Action and state are 17-dimensional: 6 + 6 arm, 2 head, 3 base velocity.

Why the compute matters strategically: 7.4 GiB cannot hold a 7-billion-parameter VLA. That rules out the OpenVLA family and forces us toward smaller models and cleverer representations rather than bigger ones. It is a constraint, but it is also what makes any result we get transferable to affordable hardware.


3. Data

Total episodes on the box298 (116,429 frames, 30 fps)
Of which marked trash/soak42
Usable256
Under the team org (xlerobot-team)252

Largest sets: xlerobot-cup-grasp-20260820-0230 (50 ep — what the current policy was trained on), xlerobot-right-pick-cup-20260826-0042 (40 ep — current gripper), xlerobot-left-pick-cup-sep3-20260903-2015 (29 ep — most recent).

All of it is cup grasping. There are zero demonstrations of pouring, shaking, stirring, presenting or handing over.

3.1 Depth is recorded but unusable by any policy

Depth exists as a side-channel: <dataset>/depth/ep<N>-<timestamp>/frame_NNNNNN.png, 640×480 16-bit PNG in millimetres, ~863 MB total.

It is not in any dataset's feature list, so no training or inference code can read it. Measured quality where it does exist: 76–86 % valid pixels, median depth ≈ 500 mm — the data itself is good. First 1–2 frames per episode are zeros (sensor warm-up); there are outliers up to 32 m that are noise and must be clipped.

Native in-dataset depth recording is implemented and waiting on a service restart.

3.2 One structural fact that shapes everything

head_1    σ = 0.0000    range [-4.308, -4.308]
head_2    σ = 0.0000    range [49.846, 49.846]

Across all 19,453 frames of the training set, the head never moved. The policy has never seen a moving camera. This is why gaze must stay frozen during manipulation and can only follow a person before and after — see §7.


4. Models and what they have actually achieved

Seven checkpoints exist locally. Trial results across 37 result files / 114 trials:

CheckpointMeasuredBatches
smolvla_20260820-1502/016000 (= last)6/50 (12.0 %, 95 % CI 4.5–24.3 %)25
a100-6000/0060000/23 (0 %, 95 % CI 0–14.8 %)3
…cup-grasp-right-b8-u20k/0200000/4 (0 %, 95 % CI 0–60.2 %)2
…right-pick-b16-3cam-u20k/020000never tested0
smolvla_20260820-1502/018000never tested0
act_20260820-0429/020000never tested0
act_20260820-0429/004000_from_hfnever tested0

Overall: 8/83 counted trials succeeded — 9.6 %, 95 % CI 4.3–18.1 %.

Two things follow that are easy to get wrong:

  1. 0/4 is not a failed model, it is an untested one. Its 95 % upper bound is 60.2 %.
  2. Four checkpoints have never touched hardware at all, including b16-3cam-u20k, which was trained on all three cameras and is therefore not comparable to the others in input form.

Zero of the 37 batches were pre-registered (a success criterion fixed before the run). By our own standard, all existing trial data is exploratory.


5. What works today

Status
Bimanual teleoperation and recording (VR and leader-arm)✅
Leader-follower interactive recording (port 8781)✅
Joint-level safety layer✅ rejected-frame rate tuned 96.6 % → 0.49 %
Control panel: jobs, datasets, checkpoints, trials, calibration (port 8770)✅
Face detection + expression + OLED expressive face + dialogue✅ as a subsystem
Live head-camera viewer with face + cup boxes, in a browser (cup_test.py, port 8795)✅ contributed by Ville — see §5.1
Head tracking of a primary face (face_follow.py)✅ not connected to the arm
Depth capture✅ side-channel only, see §3.1
Trial records: schema v2, Clopper-Pearson intervals, failure modes, funnel, pre-registration✅ new 2026-09-03/04
File drop onto the robot (port 8790)✅ new 2026-09-04

5.1 A live viewer arrived from the remote side

03-software/conversational_ai/cup_test.py (165 lines) serves the head camera in a browser with face and cup boxes drawn on it: ./run.sh cup_test.py, then open port 8795. It uses YuNet for faces — the same choice we reached independently — and EfficientDet-Lite0-int8 for cups.

Two design details worth keeping: it finds the Orbbec camera by name rather than by a /dev/videoN number, and it falls back to the shared frame in /dev/shm when another program holds the camera. Both directly avoid failure modes this project has hit before.

One caveat that matters (§6.1): it opens the camera at 1280×720, whereas the head camera is recorded — and therefore seen by the policy — at 424×240. Nine times the pixels. A detector that looks convincing in the live viewer may still fail on the frames the policy receives, which is exactly what our offline evaluation found for its default cup model.

A useful corollary: the camera evidently can deliver 1280×720, so raising the recording resolution may be a configuration change rather than new hardware. <span>UNTESTED — bandwidth, disk and retraining cost all unknown</span>


6. What does not work, or does not exist

Reliably locating the cup🔧 Was four failures; OWLv2 changed that on 2026-09-04. A zero-error configuration now exists on 36 reviewed frames, but scene generalisation is untested and the acceptance criterion needs rewriting. Earlier history: Three hand-written detectors (79 / 97.5 / 97.6 % detection, jitter to 0.9 mm) were all locking onto the robot's own arm. On 2026-09-04 four pretrained open-vocabulary detectors were evaluated offline with human review of every frame — best result 5/36 = 13.9 % on the target cup, against an acceptance bar of 90 %. See §6.1
Camera intrinsics❌ FX=FY=320 is a guess
Hand-eye extrinsics❌ never done
Depth↔RGB time alignment❌ never verified
trial head preset❌ does not exist, and everything geometric depends on it
Any policy reading depth❌ SmolVLA, ACT and every lerobot policy are RGB-only
Face subsystem connected to the arm❌
Identity recognition❌ — confirmed 2026-09-04: the face stack is YuNet, a detector. No identity anywhere
Pouring / shaking / stirring / presenting❌ no demonstrations, no policy
Keyframe capture in trials✅ implemented 2026-09-04, --dry verified

6.1 Cup detection — measured 2026-09-04, and it failed acceptance

Four pretrained detectors, run offline over recorded frames, every frame reviewed by eye:

Revised twice on 2026-09-04. First downward, after every box was cropped and identified by eye (contact-sheet thumbnails had made a VR controller look like a cup). Then substantially upward, after switching to OWLv2.

The first pass — four detectors, all failing

DetectorBoxesActually a cupThe target cupLatency
EfficientDet-Lite0 int80/36——66 ms CPU
EfficientDet-Lite27/36not itemised—411 ms CPU
OWL-ViT base-patch3216/3610 (62.5 %)5225 ms GPU
Grounding DINO tiny27/36112657 ms GPU

OWL-ViT's false positives were a VR controller left on the table (5 boxes); Grounding DINO's were cable-grommet holes in the table, the mouse, the gripper and cabling (26 of 27).

A claim was retracted here. An earlier version said *"recognition is fine, it just picks the
wrong cup"*. That was too generous: 6 of OWL-ViT's 16 boxes and 26 of Grounding DINO's 27 were
not cups at all. Precision and disambiguation were both broken.

The second pass — OWLv2 changes the picture

Same 36 frames, same queries, same selection rule; only the model changed.

ConfigurationBoxesTarget cupOther cupNot a cupVerdict
OWL-ViT @0.051655662.5 % are cups
OWLv2 @0.0524183387.5 % are cups
OWLv2 @0.30191630100 % are cups
OWLv2 @0.30 + background difference @50 %151500first zero-error configuration

The VR controller and the table holes produce zero boxes under OWLv2.

And thresholding went from useless to decisive, which is the structural change:

Target-cup scoresNon-cup scoresSeparable by threshold?
OWL-ViT0.051 – 0.1250.058 – 0.213❌ the false positives score higher
OWLv20.18 – 0.710.09 – 0.25✅ 0.30 cuts cleanly

What this costs, and what is still unknown

Latency1623 ms/frame — 7× OWL-ViT, and slower than a full SmolVLA inference (1329 ms). Per-frame online use is impossible. Viable as a staged call: "stop, look once, lock the target"
Per-frame detection rate41.7 % (15 boxes from 36 frames)
Scene generalisation❌ not tested at all. 36 frames, one room, one lighting condition
The background referenceA cross-episode pixel median, not a real empty-table photograph — and it is contaminated by motion blur from the arm. So the layer's measured performance is a lower bound
Acceptance cannot yet be declared passed, and the reason is worth recording.
The GOAL fixed "≥30 frames reviewed by eye, ≥90 % correct" — but "correct" turns out to have
at least three readings: precision 100 % (15/15 boxes are the target cup),
per-frame rate 41.7 %, and "found when a cup is actually visible" — **which nobody has
measured**, because it was never counted how many of the 36 frames contain a visible,
unoccluded cup at all. The criterion did not distinguish these. **It needs rewriting into two
separately pre-registered numbers, and the second one re-measured per episode rather than
per frame** — that is the metric that matches the actual use ("look once before reaching").

One operational prerequisite that will fail silently if forgotten

Background subtraction requires the reference image and the runtime camera pose to match. The head is locked during trials, so within a session this holds. But changing the head preset invalidates the reference, and nothing will warn you. This must go into the operating procedure.

6.2 Two things found while re-checking, both cheap to act on

  1. A measurement discipline caught a second mistake of the same kind. The rule adopted was: every filter must report two numbers — false positives removed AND true positives lost. Re-scoring earlier interventions under that rule overturned one verdict:

| Intervention | False positives | Target cups | Net | Previously called | |---|---|---|---|---| | OWL-ViT, drop "a drinking glass" | 6 → 7 (worse) | 5 → 1 | loss | correctly, a loss | | Grounding DINO, drop "a drinking glass" | 26 → 1 | 1 → 1 (no loss) | clear gain | wrongly, "over-filtered" | | OWL-ViT → OWLv2 | 6 → 3 | 5 → 18 | double win | — | | OWLv2 threshold 0.05 → 0.30 | 3 → 0 | 18 → 16 | worthwhile | — |

The Grounding DINO run had been dismissed because only the box count was reported — it dropped to 2/36 and that looked like a collapse. Counting both numbers shows it removed 25 false positives at zero cost. Reporting one number instead of two produced a wrong conclusion twice in one investigation.

  1. The cascade's layers need reordering, because layer 1 turned out to do layer 2's job.

| Layer | Was meant to remove | Left to remove after OWLv2 | Its value now | |---|---|---|---| | ① Semantic (open-vocabulary + corrected distractors) | arm, mouse, cables | — | non-cup false positives are zero | | ② Depth plane residual | flat things (the table holes) | 0 — the holes no longer produce boxes | ≈0 as a filter, but still the only source of height/distance for grasping | | ③ Empty-table background difference | permanent scene objects | 3 (the other plastic cup) | the only layer still removing errors |

New order: ① → ③ → ②. This is not a rejection of depth — its role changed from filtering to supplying 3D position, and that still needs intrinsics and depth/RGB alignment, which are on-site work. It should no longer block validating layer ③.

  1. Background subtraction at 40 % is free. Removing boxes whose changed-pixel fraction is below 40 % kills 2 of the 3 "other cup" boxes and costs zero target cups. At 50 % it removes all 3 and costs 1 target cup.

7. The design decisions that are settled, and why

7.1 Gaze is phase-based, and this was forced by data

#StageHeadArm
1Greet, conversefollows the facestill
2Make the drinklockedpolicy
3Rotate the glass to show itfollows the facepolicy
4Push it toward the customerfollows the facepolicy

Two independent constraints force the middle row: the policy has never seen a moving head (§3.2), and the depth camera's extrinsics are only valid for one head pose. This was not chosen; it is what survived elimination.

Stage 1 is the cheapest demonstrable win in the project — the arm is still, so it conflicts with nothing, and no commercial product does it.

7.2 First version pushes the glass; it does not hand it over

Handing a glass into someone's hand is a human-robot handover problem whose core difficulty is when to release. Too early and it drops; too late and the person cannot take it. Humans use touch. We have no force sensing (§2). Pushing the glass along the table keeps support under it the whole time and needs no release decision.

7.3 Learned motion, not recorded trajectories

This is the project's one real technical bet, and it is unproven — see §4.

7.4 Every claim carries its evidence level

A worked example from 2026-09-04. Several documents in this repo — including an earlier draft of this one — stated "SmolVLA runs at 146 ms/frame (6.9 Hz)". Re-measuring produced 1329 ms, nine times larger. Tracing it back: the 146 ms figure was a real measurement from 2026-08-21, but of ACT, not SmolVLA; the model name was substituted somewhere in the chain of hand-off documents. Re-measuring both on the same machine confirmed it: ACT 129.5 ms, SmolVLA 1329.2 ms.

A second, framework-level error travelled with it: "6.9 Hz cannot keep up with 30 Hz training" is not a meaningful statement about a chunked policy. SmolVLA plans 50 action steps per inference, which is 1667 ms of motion at 30 Hz. The right question is whether one inference fits in 1667 ms — it does, at 1329 ms, with 337 ms to spare.

Both errors survived multiple documents because nobody re-measured. That is the failure mode this rule exists to catch.

measured / not measured / tried but blocked. Adopted after a review found untested inferences being reported as facts. It applies to other people's papers too: Oat-VLA reports 59 % vs 41 % on real robots; recomputing from its own counts gives 29/49 vs 20/49, overlapping intervals, Fisher p = 0.106 — not significant. We cite its token reduction, not its success rate.


8. Known architectural conflicts

ConflictStatus
1face_follow.py and the policy both own the left serial bus, and the safety layer overwrites the head every frame. They cannot run simultaneouslyResolution chosen: head control stays inside the policy process, only the target source changes per phase. Not implemented
2The safety layer has no concept of a person. Human safety today = a fixed geometric envelope + someone holding the e-stop. "Reach toward the customer" is the exact motion that envelope is shaped to preventNeeds a parallel design. Not started
3The per-frame step cap was derived from grasping data (wrist_roll ≤ 10.4°/frame) and it silently slows motion rather than rejecting it. A batch of presenting demonstrations could be recorded flattened, invisiblyMitigation: measure the motion once before recording. Pending on-site
4Three face detectors would exist if Ville's is added (Haar in scripts/, YuNet in conversational_ai/, his). Convergence neededOpen

One factual correction found on 2026-09-04: policy_safety.py's own note asserts the training set carries head_motor_2 = 99.649, and concludes from that arithmetic that the head sat 564 counts outside the calibrated range — which is why the absolute clamp skips the head entirely. The data says 49.846, which recomputes to inside the range by 2 counts. The premise is wrong. Harmless while the head is pinned; not harmless once the head moves.


9. Market position

Verified from vendor pages and named sources (full analysis with images: 04-reports/market-analysis/, not published publicly):

SystemFormThroughputPrice
Makr Shakr Toni2 arms, 158 bottles80+/hr (official)—
Makr Shakr Veloce2 arms, speed-oriented240/hr—
Cecilia.aiScreen persona + pumps, voice dialogue120/hr$45,000 or $2,000/mo
Yanu (Tartu, Estonia)Arm inside an enclosed pod, dialogue + payments100/hr—
MIXO / TendedBar / BotristaPumps, carousels—industry band $30k–$100k+

The gap, stated narrowly enough to survive checking: conversational bartender robots exist (Cecilia, Yanu) and expressive robot arms exist (Makr Shakr), but no product puts both in the same body — and in all of them the arm motion is pre-programmed. Their sourced description: "The recipe database contains pre-programmed ratios, sequences, and pour volumes", while machine learning "tracks popular combinations, personalizes suggestions". The AI is in the recommendation layer, not the motion layer.


10. People and where they are

RoleLocationReaches the robot how
Project leadChinaTailscale (100.107.145.111)
Face recognition & interaction (Ville)Taiwan (exchange)Not yet on the network — an invite would let him measure on the real board
On-site data collection & hardwareWith the robotPhysically

Only one person can touch the machine. Everything else is remote, which is why the on-site brief (BRIEF-ONSITE-HARDWARE.md) is written around "anything you don't measure while you are there, nobody can measure later".


11. What happens next

The open decision — this one blocks everything downstream

Since detection failed on disambiguation rather than recognition, four routes exist. Recommendation from the analysis: (a).

RouteAssessment
aUse depth to disambiguateRecommended — depth is the only signal that is insensitive to appearance, and we already have 863 MB of it. Blocked on camera intrinsics and depth/RGB alignment, which are on-site work
bRaise head-camera resolution424×240 is genuinely small, but this does not address ranking
cBetter language promptsCheap to try; unlikely to fix ranking
dFiducial markers on the cupsReliable fallback, but concedes the generalisation claim

On-site (needs the robot)

  1. Save the trial head preset — 1 minute, everything geometric depends on it
  2. Trial the four never-tested checkpoints, with pre-registration, 10 runs each
  3. Camera intrinsics → hand-eye extrinsics → depth/RGB time alignment
  4. Restart 8781, record 3–5 episodes of native-format depth for verification
  5. Measure joint speed of the rotate-the-glass motion before recording that skill

Remote (parallel, no hardware)

The ordering principle

Nothing downstream matters until the cup can be reliably located. Pouring, presenting and handing over all begin with a successful grasp, and grasping currently succeeds 8 times in 83.


12. The three ways this most plausibly fails

RiskFirst mitigation
1Grasping never gets good enough. 8/83 today. Detection failed four times, then OWLv2 produced the first zero-error configuration (15/15 on reviewed frames). The risk eased but did not close: one room, one lighting, 36 frames, and per-episode hit rate never measuredWiden to ~150 frames across batches to rule out luck; then re-measure per episode
2Demonstrations cannot teach "elegant". Pour rate, stop timing and style exist only in the demonstrations, and we have none. Better perception does not tell a model how much to pourMeasure the motion first (§8 item 3), then record deliberately
3Compute runs out — now quantified. One SmolVLA inference is 1329 ms and plans 50 steps = 1667 ms of motion at 30 Hz — 80 % of budget used, 337 ms margin. Any per-frame detector eats that marginMeasured 2026-09-04. Grounding DINO tiny alone is 2657 ms, so online use is out

All three have a first step that is remote, cheap, and can be done before the next batch of demonstrations is recorded.


Companion documents: BRIEF-FACE-INTERACTION.md · BRIEF-ONSITE-HARDWARE.md · HANDOFF-REMOTE-2026-09-04.md · TOMORROW-2026-09-04.md · three GOAL files · 04-reports/market-analysis/ · STATUS.md (machine-generated, always current).