The market has already solved making the drink. We are solving something else: making the drink worth watching.
To be clear: these products are not bad. They have a different objective — throughput and consistency — and they hit it.
| System | Form | Throughput | Price |
|---|---|---|---|
| Makr Shakr Toni | Two arms (KUKA), 158 bottles | 80+ /hr (official page) | — |
| Makr Shakr Veloce | Two arms, built for speed | 240 /hr | — |
| Cecilia.ai | Screen persona + pumps · voice dialogue | 120 /hr | $45,000 or $2,000/mo |
| Yanu | Arm inside an enclosed pod · dialogue + payments | 100 /hr | — |
| MIXO 8 / 20 | Peristaltic pumps | — | — |
| TendedBar / Backbar One | Carousel + valves | built for high throughput | — |
The installation is designed around the machine. Bottles inverted on a fixed rail, glasses in a fixed jig, motion along a fixed path. Under those conditions a pre-programmed trajectory is not a compromise — it is the right engineering choice, and it is exactly why their demo footage looks so fluid. That is their strength, not their weakness.
This is the hinge of the whole argument, and it is quotable.
“The recipe database contains pre-programmed ratios, sequences, and pour volumes.”
“machine learning tracks popular combinations, personalizes suggestions, and adjusts pours based on historical data.”
The intelligence decides what you should drink. The arm still replays a recording.
Robots that talk exist. Robots that perform exist. Nothing puts both in one body.
And one property cuts across all of them: talkative or not, every arm on this chart replays a pre-programmed trajectory. The AI is in recommendation, or in dialogue — never in the motion. That is the only real technical bet this project makes.
If all you want is the drink, a pre-mixed can is cheaper. People sit at a bar because there is a person there, a performance, an exchange.
240 drinks/hour. Right for stadiums, cruise ships, airports — places with a queue.
Looking up at you, turning the glass so you see the colour, sliding it across to you.
Where there is no queue — craft bars, showrooms, events — that stretch is worth more than speed.
So we are not racing 240 drinks/hour. We are building the one that lifts its head.
Watch the orange dot: gaze appears twice, at the beginning and the end, and is deliberately frozen in between. That is not a preference — slide 10 gives the evidence.
"Knowing where the customer is" is the shared prerequisite of the first and last stages. No existing product needs it, because their customer walks to a collection point. That collection point is the thing we are removing.
These sound like variants of one thing. They differ by an order of magnitude.
The glass never leaves the surface.
Why it works: supported throughout; failure modes are gentle (off-target, tipped); no need to decide when to let go.
A real human-robot handover. The core difficulty is release timing: too early and it drops, too late and the person cannot take it.
Why it is blocked: humans judge this by touch, and we have no force or torque sensing — the observation is 17 joint positions.
UNVERIFIED Without force sensing, only two routes remain to detect "they have taken it", and neither has been tried: read it from the wrist camera (heavy occlusion), or use servo load — the servos report it, but it is not currently in the observation at all, so the recording pipeline would have to change first.
This must be decided before demonstrations are recorded. The two motions need different data; choosing wrong means recording it all again.
Four frames pulled from our recorded demonstrations, straight from the robot's head camera at its real resolution (424×240).
xlerobot-team/xlerobot-cup-grasp-20260820-0230,
50 episodes / 19,453 frames / 30 fps. Frames extracted with ffmpeg; no cropping or colour grading.A vision-language-action model trained from human demonstrations. A different glass or a moved bottle is handled by generalisation, not by re-teaching. The cost is that it is harder, and has to be proven.
Face detection, expression recognition, an OLED expressive face and voice — already running as a subsystem. Not decoration bolted on at the end: it is the input that decides where to look.
Rotating to show, handing toward a person. These do not exist in current products because their customers collect their own drink.
Rather than feeding the model a 3D coordinate we computed ourselves, we group the visual tokens by object and let the model decide what deserves attention. With a bottle, a glass and a person in frame at once, that difference is decisive.
MEASURED — AND IT CORRECTED US There is precedent on a real VLA (Oat-VLA: tokens 256→16, twice as fast to converge on LIBERO). But measured on our board on 2026-09-04, the speed argument does not hold: all three cameras together account for only 177.7 ms — 13 % of total inference. Removing every visual token could save at most that, and the segmentation model is itself a ViT. The speed rationale is dead. The structural-prior rationale survives, resting on one non-significant number.
MEASURED Two arms, a moving head, three cameras, a mobile base. This board cannot hold a 7-billion-parameter model — our technical choices are genuinely constrained by compute. That is not a weakness: it forces us toward smarter rather than bigger, like the 94 % token reduction on the previous slide.
A product decision that was ultimately settled by one line of numbers.
In every demonstration, the head never moved once. The model has never seen footage from a moving camera. Turning the head toward a customer mid-task puts the input outside everything it has ever seen — it would not be slightly worse, it would be a situation it has never encountered.
And a second, independent reason: the depth camera is mounted on that moving head, so its calibration is valid for one head pose only. Move the head and the geometry is void.
Two independent constraints point at the same answer, so "look at the person while
greeting and handing over, freeze while working" was not chosen — it is what survived
elimination.
Good product decisions often look like this: not selected, but left standing.
This slide is not usually in a pitch. It is here because it is our real difference.
MEASURED ran it, have the output UNVERIFIED did not
An untested inference may not be reported as a tested fact. This rule exists because we broke it.
0/4 and 0/23 look identical in a table, but the 95 % upper bound on the true success rate differs by a factor of four. See the chart below.
We built three cup detectors. Detection 97.6 %, jitter 6.4 mm — every number excellent. Looking at the images showed all three had locked onto the robot's own arm. Three times.
Oat-VLA reports 59 % vs 41 % on real robots. Recomputing from its own counts: 29/49 vs 20/49, overlapping intervals, Fisher p = 0.106 — not significant. So we cite its token reduction and not its success rate.
The slides above are where we are going. This one is where we stand, including the parts that do not look good.
We put 8/83 in the pitch because that is how this project works. A robotics project that claims success has usually just not counted carefully. We counted, and we published the confidence interval too.
| What | Why it is next | Needs the robot | |
|---|---|---|---|
| 1 | Locate the cup reliably (open-vocabulary detection + human review of the images) | Every later skill starts here | no |
| 2 | Trial-record discipline: criteria fixed in advance, failure modes, intervals | Otherwise "improved" and "got lucky" are indistinguishable | no |
| 3 | Camera intrinsics + hand-eye + a saved head preset | Prerequisite for any 3D reasoning | yes |
| 4 | Wire up the greeting stage (arm still, gaze follows) | Conflicts with nothing; first demonstrable piece | yes |
| 5 | Record pouring / showing / handing demonstrations | These skills can only come from demonstrations | yes |
| 6 | A safety policy for a person in the workspace (designed separately) | Required before reaching toward anyone | no → yes |
Item 4 deserves a separate mention: during greeting the arm is still, so it conflicts with nothing, and it is the first piece of the whole story that can be shown to someone — a robot that looks up at you, talks, and changes expression. That alone is already outside what the existing products do.
If this project does not work, it will probably be one of these. They are written down so they can be watched.
8/83 today. If "locate the cup reliably" cannot be solved, pouring and presenting are moot. This is the primary risk, and the reason it is item 1 on the plan.
Pour rate and stop timing live only in the demonstrations. No amount of perception tells a model how much to pour. And we have recorded none of these yet.
One SmolVLA inference takes 1329 ms and plans 50 action steps — at 30 Hz that is 1667 ms of motion, so 80 % of the budget is used, leaving 337 ms. Add a detector that must run every frame and the margin is gone.
All three have a first step that is remote, cheap, and costs no robot time: run detection offline and look at the images (one); measure the joint speed of the presenting motion before recording it (two); measure the net effect of token reduction (three). All three can be settled before the next batch of demonstrations is recorded.
Existing bartender robots treat the customer as someone who collects a drink.
We are building the one that treats them as someone who is watching.
That difference decides every technical choice: why the motion is learned rather than recorded, why gaze switches by stage, why version one slides the glass instead of placing it in a hand, and why 8/83 is on a slide in this deck.
04-reports/market-analysis/ and are not published publicly.