The market has solved making the drink. We are solving something else: making the drink worth watching — and in this version the arm's motion is learned, not recorded.
Makr Shakr 240 drinks/hour, Cecilia.ai $45,000, Yanu in a sealed pod: optimised for throughput and consistency, and they hit it. The installation is built around the machine — glasses in a jig, bottles on a rail, motion along a fixed path.
“The recipe database contains pre-programmed ratios, sequences, and pour volumes.”
“machine learning tracks popular combinations, personalizes suggestions, and adjusts pours based on historical data.”
Robots that talk exist; arms that move exist; nothing puts both in one body, and every arm replays a recording. Move the glass 5 cm and the same trajectory misses.
Only one stage runs at a time; every other button is disabled while it does. The policy process and the scripted stages both take both serial buses.
Between stages the arms are holding something, so every stage keeps torque on exit. Drop torque and the drink goes on the floor.
Deliberately. A person looks between every stage — that is the point of one manual button per stage. Every live stage is recorded from the head camera.

A DQ store with a sharpa humanoid behind glass. It makes one product — the Blizzard. You cannot walk in for it: you book a slot in a mini-programme, and slots for the next week open on Friday and Saturday at 10:00. The place was packed while we were there.
One flavour, 55 scripted steps, no learned motion, no gaze — and it still draws a queue and pays for the floor it stands on. That is the commercial case for the 40 seconds, already proven by someone else. It is also the ceiling of that approach: the show is the same every time, and the machine never looks up. Ours looks at the customer, and its arm motion is learned rather than replayed.
Evidence level: our own photos and video on site, 2026-09-12 — the strip above is the raw material. Signage and screen text are transcribed from those frames. "Packed", "popular" and "profitable" are our impression on the day; no footfall or revenue figures were taken, and the operator publishes none.
A bar has never only sold liquid. People sit at a bar because there is a person, a performance, an exchange. Existing systems optimise speed for places with a queue; we optimise the 40 seconds you are watched — craft bars, events, showrooms.
Warm, quick and a little silly. One or two sentences, already reaching for the glass while talking, and the line while fetching ties their drink to them.
Teases the drink, never the customer. Never says a price.
Listens to the small talk — just will not follow you up into the clouds. When the customer turns wistful it pulls the conversation back down to the floor: "that is not a philosophical question."
Dry, practical, paying attention the whole time. The opposite temperature, the same body.
A bartender is a text file here. Swap the file and you have swapped the bartender; the head wears the outfit that goes with it. That is how one machine fits a craft bar and a trade-show stand without rebuilding anything.
"Looking at the customer" needs a gaze that turns and a face that changes. Gaze is on the person while greeting and handing over, locked while making.
Across 54 demonstrations the two head joints have std 0.0000 — the model has never seen a moving viewpoint. The depth camera rides on the head; move it and the calibration is gone.
Pan range 342–3754 is set by the depth-camera cable; the standard view cup-grasp-v0 = pan 2003 / tilt 2617, aimed and confirmed by image before every run.
The right-arm model scored 0/14 on 09-08. Diagnosis: cup-position coverage in the 40 demos was 16× unbalanced (16 episodes in the densest band, 1 in the sparsest); the model had learned tempo, not vision — closure timing correlated with cup position at only r = 0.16–0.26.

"It was about to push the cup away, adjusted before it did, and grasped it. Very smart."
"At first it pushed the cup outward a little, then corrected immediately and lifted it."
That is what a closed loop looks like: miss → see → correct. A flawless run proves nothing; a correction proves the policy is reading the image.
49 live runs on record, across two arms. Left arm, 2026-09-09: 28 runs over four different objects, 15/20 at its best configuration. Right arm, 2026-09-11: 6/6 on the cup it was trained on — then, same checkpoint, no retraining, 3/8 and 1/5 on cups and positions it was never trained for, spread over eight bands from −40° to +14°.
Those three numbers stay apart instead of averaging into one flattering figure — the counting rule on slide 8 doing its job. And apart, they say the useful thing: solid on what it knows, degrading rather than vanishing on what it does not, and we can point at the edge. Anyone can show you a robot that works once.
Slide the cup 10 cm in front of the audience and let it find the cup again. The silent question all along is "is this a recording?" — this answers it, and costs nothing.
The training set has one cup per frame and one sentence. With two cups on the table it reaches between them (three times on 09-11); adding "yellow" to the instruction perturbs the trajectory but not the target. SmolVLA has 450 M parameters on a 7.4 GB board; teaching it to select by word means re-recording and a night of training, with no precedent that it would learn.

The rule changed from "paint out the cup of colour X" to "keep one, paint out everything else". A second distractor on the table should not need a config change.
On a 25-frame evaluation set the colour thresholds we had measured in daylight on 09-11 detected 1 of 25. At night the table and the cups both fall to saturation 10–30 and colour stops being a separable feature at all. Rebuilt around the table region plus brightness relative to the table: 15 of 25.
The gripper cannot be excluded by colour either — the cream cup and the orange bracket share one hue band (H 8–30), so excluding orange kills the cup. It is excluded by position: both arms always enter from the bottom corners.
And the "empty table" plate we meant to paste over distractors turned out not to be empty — it had a green cup in it. Pasting it would have inserted a cup that is not there. Median fill instead: no plate, no shape, adapts to the light.
Selection stays out of the model. Everything on the table except the chosen cup is painted out of the head image and filled with the median colour around it; the policy sees the single-cup scene it was trained on.
3.3 ms per frame (424×240, CPU), 3% of the control period. Same checkpoint; a new instruction only changes the mask colour.
The small model's generalisation is spent on what it is good at — grasping. "Which cup" is a 3 ms classical-vision problem, not worth a night of GPU.
measured25 real head frames, day and night: the 09-11 colour rule detects 1/25, the rebuilt detector 15/25. Two new channels shipped — --mask-keep (keep one cup, paint out the rest) and --mask-box (paint out a rectangle). Runs in the pipeline with per-order switching, 3 ms/frame. Eight exploratory trials on the left-arm model went for the blue cup whichever cup was masked — that model carries its own bias and is not a valid testbed. not measuredselection accuracy — needs a preregistered 5 + 5 on the right-arm model (mask blue → grasp yellow, mask yellow → grasp blue) with the criterion "closure pan on the target side". The A/B on 2026-09-12 does not count: there was no target cup in the grasp area, so both rounds were held=0 by construction and the 2% safety-layer difference is single-trial noise. It needs a target and a distractor on the table, and someone on site to place them.
The team is in China, the robot in Finland. Every trial on 2026-09-11 — launch, judging, logging — was run remotely: the panel starts the job, the person on site places the cup and presses confirm, results are archived automatically.



Three sites, one robot. The arm is in Helsinki; the people who launch the runs, judge them and write them up are 7,000 km away in Hangzhou and Taiwan. Nothing on this deck needed anyone to fly.
Teleoperation → demonstration recording → dataset → cloud training → autonomous execution → human takeover. Every link in it ran on 2026-09-11.
VR headset teleop, leader–follower recording (8781), the 8770 panel for trials and control. Tailscale, travels with the robot.
One remote team can watch several machines; on site you need one person to place cups and hold the e-stop.
Why bother. Not modesty — this is what lets us say the interesting things and be believed. The three numbers on slide 5 are what the discipline produces: 6/6, 3/8, 1/5, kept apart rather than averaged into one figure that would flatter us and tell you nothing.
Two arms, a moving head, three cameras, a mobile base. The board cannot hold a 7-billion-parameter model — so we run the 450 M SmolVLA and hand problems like "which cup" to 3 ms of classical vision (slide 6). The constraint forces a smarter design, not a bigger model.
| Cannot do today | Evidence | Next step |
|---|---|---|
| Two cups on the table: it reaches between them | three times on 09-11, traces on file | preregistered 5+5 with the mask; or 50 episodes with colour instructions |
| The left-arm model grasps only the blue cup on the right, whatever it sees | consistent across four conditions plus 8 masked trials on 09-11 | run colour experiments on the right-arm model; record two-cup data for the left |
| The right arm half-recognises cups of other colours | blue 1/2, pink 0/1, green can 0/1: reaches the spot, does not close | as above, or the mask to restore the training-cup scene |
| Selecting a cup by language | A/B with "yellow": target unchanged | as above |
| Pouring is scripted; no liquid on stage | pour_cup.py, one hardware run 09-09 | a POUR state in the state machine; empty cups for the demo |
| A different table | 3 cm of table height voids the policy | venue-move procedure written: reproduce shoulder-to-table to ±1 cm |
| Left shoulder servo overheats | 50 °C+, 6–7° under-torque | replace, then full recalibration |
A promise, not a wish list. Every line below carries the condition that decides whether it passed, and every result lands in the public report space with its date and its k/n — including the ones that fail. You will not have to take our word for any of it.
✓ Already done: greet → grasp → hand over in one continuous run, one process, nobody touching the robot, on uncut video.
50 episodes, half yellow and half blue, the two positions swapped, trained together with the single-cup set.Done when: the dataset is public and the coverage CV across the five cup-position bands is reported.
Criterion, trial count and void rule fixed before the run, then judged under the slide-8 rules.Done when: the k/n is published — 6/6 or 3/10, it goes on the same page.
Pouring moves inside the pipeline: a POUR state and a pour() effector. Today coke, fanta, sprite and water are four names for one physical skill; this ends that.Done when: at least two drinks of two or more ingredients, poured in order to a recipe, run end to end on video.
Ten consecutive runs of the same configuration with no hand in the loop, and the hand-over direction taken from the measured face bearing rather than a fixed coordinate.Done when: the ten-run k/n is published under the same rules.
The 10-step arrival procedure at a new venue, two hours from boxes to a run; one cocktail event or bar week agreed and attended; feedback collected on the night.Done when: the event happened, the feedback is written up and published whatever it says, and a named venue has stated in writing what would put the robot on their floor.
Not a success rate. The counting rules on slide 8 do not move to flatter us: k/n per configuration, never pooled, void trials named and excluded, evidence level on every claim. If the colour-instruction round comes back 3/10, that is what the report will say, on the same page as the 6/6.
That difference decides every technical choice: why motion is learned rather than recorded, why gaze switches by stage, why cup selection sits in front of the model, why every number carries its source.
