← reports index
XLeRobot · 2026-09-04

Not a faster drink machine.
A bartender that looks at you.

The market has already solved making the drink. We are solving something else: making the drink worth watching.

240drinks / hour
Makr Shakr Veloce
$45,000Cecilia.ai list price
industry band $30k–$100k+
0products that put an expressive body
and awareness of the customer together
The first two are sourced — see slide 2. The third needs stating precisely: conversational bartender robots do exist (Cecilia.ai, Yanu), and expressive robot arms exist (Makr Shakr) — but no product puts both in the same body. Slide 4 breaks that down.
01 · The market

What they already do well

To be clear: these products are not bad. They have a different objective — throughput and consistency — and they hit it.

SystemFormThroughputPrice
Makr Shakr ToniTwo arms (KUKA), 158 bottles80+ /hr (official page)—
Makr Shakr VeloceTwo arms, built for speed240 /hr—
Cecilia.aiScreen persona + pumps · voice dialogue120 /hr$45,000 or $2,000/mo
YanuArm inside an enclosed pod · dialogue + payments100 /hr—
MIXO 8 / 20Peristaltic pumps——
TendedBar / Backbar OneCarousel + valvesbuilt for high throughput—
Why they can be this fast

The installation is designed around the machine. Bottles inverted on a fixed rail, glasses in a fixed jig, motion along a fixed path. Under those conditions a pre-programmed trajectory is not a compromise — it is the right engineering choice, and it is exactly why their demo footage looks so fluid. That is their strength, not their weakness.

Sources: makrshakr.com/toni (official) · cecilia.ai · yanu.ai · sedonatec.com · designboom
Toni's 80+ is the vendor's own figure; Veloce's 240 is press-reported, with a secondary source stating 250 — vendor numbers take precedence, press numbers are labelled.
02 · The gap

But which layer is their AI in?

This is the hinge of the whole argument, and it is quotable.

Motion layer

Pre-programmed

“The recipe database contains pre-programmed ratios, sequences, and pour volumes.”

AI layer

Recommendation

“machine learning tracks popular combinations, personalizes suggestions, and adjusts pours based on historical data.”

The intelligence decides what you should drink. The arm still replays a recording.

PRE-PROGRAMMED TRAJECTORY glass in the jig → fixed path → success glass moved 5 cm → same path → misses LEARNED POLICY (US) look → compute action → success glass moved 5 cm → look → should also work ←unproven
Quotations: sedonatec.com. The right-hand box is labelled unproven deliberately — our own grasp rate is 8/83, see slide 12. This deck does not present untested things as tested.
02b · Positioning

Where the actual gap is

Robots that talk exist. Robots that perform exist. Nothing puts both in one body.

a body capable of expressive motion → aware of the person → MIXO · TendedBar pumps and valves; no body Makr Shakr two arms · 158 bottles · 80+/hr no customer perception marketed Cecilia.ai · Yanu voice, dialogue, payments a screen persona, or an arm sealed inside a pod XLeRobot two arms + moving head + 3 cameras face, expression, dialogue motion is learned, not scripted ← this quadrant is empty

And one property cuts across all of them: talkative or not, every arm on this chart replays a pre-programmed trajectory. The AI is in recommendation, or in dialogue — never in the motion. That is the only real technical bet this project makes.

Placements are based on each vendor's own pages and public reporting. "No customer perception marketed" refers to the full text of makrshakr.com/toni — it does not prove the product lacks the capability, only that the vendor does not sell it as one.
03 · The thesis

A bar has never only sold liquid

If all you want is the drink, a pre-mixed can is cheaper. People sit at a bar because there is a person there, a performance, an exchange.

What the market optimises

Speed of service

240 drinks/hour. Right for stadiums, cruise ships, airports — places with a queue.

What nobody optimises

The 40 seconds you are watched

Looking up at you, turning the glass so you see the colour, sliding it across to you.

Our bet

The interaction is the product

Where there is no queue — craft bars, showrooms, events — that stretch is worth more than speed.

So we are not racing 240 drinks/hour. We are building the one that lifts its head.

04 · The product

One service is four stages

Watch the orange dot: gaze appears twice, at the beginning and the end, and is deliberately frozen in between. That is not a preference — slide 10 gives the evidence.

GAZEARM 01 Greet · converse turns to find you, takes the order, the face responds 02 Make the drink grasp · pour · shake · stir gaze must be locked 03 Show the drink rotates the glass so you can see colour and layers 04 Hand it over toward where you actually are — not a fixed point still policy-driven ————————————————→ stage 1 has a still arm — nothing to conflict with, and it can be demonstrated today

"Knowing where the customer is" is the shared prerequisite of the first and last stages. No existing product needs it, because their customer walks to a collection point. That collection point is the thing we are removing.

05 · Handing over

Slide it across, or place it in a hand?

These sound like variants of one thing. They differ by an order of magnitude.

Option A · version one

Rotate, then slide it across

The glass never leaves the surface.

Why it works: supported throughout; failure modes are gentle (off-target, tipped); no need to decide when to let go.

Option B · later

Rotate, lift, place it in a hand

A real human-robot handover. The core difficulty is release timing: too early and it drops, too late and the person cannot take it.

Why it is blocked: humans judge this by touch, and we have no force or torque sensing — the observation is 17 joint positions.

UNVERIFIED Without force sensing, only two routes remain to detect "they have taken it", and neither has been tried: read it from the wrist camera (heavy occlusion), or use servo load — the servos report it, but it is not currently in the observation at all, so the recording pipeline would have to change first.

This must be decided before demonstrations are recorded. The two motions need different data; choosing wrong means recording it all again.

06 · What it sees

Not a render — its actual field of view

Four frames pulled from our recorded demonstrations, straight from the robot's head camera at its real resolution (424×240).

Four head-camera frames: cup on the table, arm approaching, gripper in position, closing
Head camera · four moments from one episode: cup on the table → arm approaches → gripper in position → closing. This is the resolution the policy actually receives.
Two right-wrist camera frames
Right wrist camera (640×480) · the red part is our own 3D-printed compliant fin-ray gripper. Left: approaching. Right: the cup enters view. Three cameras — one on the head with depth, one on each wrist.
Source: dataset xlerobot-team/xlerobot-cup-grasp-20260820-0230, 50 episodes / 19,453 frames / 30 fps. Frames extracted with ffmpeg; no cropping or colour grading.
07 · Approach

Three things we do differently

01

Motion that is learned,
not recorded

A vision-language-action model trained from human demonstrations. A different glass or a moved bottle is handled by generalisation, not by re-teaching. The cost is that it is harder, and has to be proven.

02

Emotion and dialogue
as first-class inputs

Face detection, expression recognition, an OLED expressive face and voice — already running as a subsystem. Not decoration bolted on at the end: it is the input that decides where to look.

03

Motion aimed at an audience

Rotating to show, handing toward a person. These do not exist in current products because their customers collect their own drink.

The key technical choice: group the image by object before the model sees it

Rather than feeding the model a 3D coordinate we computed ourselves, we group the visual tokens by object and let the model decide what deserves attention. With a bottle, a glass and a person in frame at once, that difference is decisive.

TODAY · 256 VISUAL TOKENS OBJECT-CENTRIC · 16 93.75% fewer — but measured, visual tokens are only 13% of the cost

MEASURED — AND IT CORRECTED US There is precedent on a real VLA (Oat-VLA: tokens 256→16, twice as fast to converge on LIBERO). But measured on our board on 2026-09-04, the speed argument does not hold: all three cameras together account for only 177.7 ms — 13 % of total inference. Removing every visual token could save at most that, and the segmentation model is itself a ViT. The speed rationale is dead. The structural-prior rationale survives, resting on one non-significant number.

08 · Hardware

It runs on a board that fits in your hand

7.4 GBJetson Orin Nano
all the memory, CPU and GPU sharing it
865 MBpolicy weights
inference is local, nothing leaves the box
$30k–100kcommercial systems
our bill of materials is an order of magnitude lower
Robot head: spherical shell with a large eye, front and isometric views
The head · near-front and isometric. Our own design: spherical shell, one large eye, pan/tilt, carrying both the depth camera and the expressive display. The camera moves — which is where the constraint on slide 11 comes from.

MEASURED Two arms, a moving head, three cameras, a mobile base. This board cannot hold a 7-billion-parameter model — our technical choices are genuinely constrained by compute. That is not a weakness: it forces us toward smarter rather than bigger, like the 94 % token reduction on the previous slide.

09 · A worked example

Why gaze must freeze while it works

A product decision that was ultimately settled by one line of numbers.

50 DEMONSTRATION EPISODES · 19,453 FRAMES · BOTH HEAD DEGREES OF FREEDOM
head_1   σ = 0.0000   range [-4.308, -4.308]
head_2   σ = 0.0000   range [49.846, 49.846]

In every demonstration, the head never moved once. The model has never seen footage from a moving camera. Turning the head toward a customer mid-task puts the input outside everything it has ever seen — it would not be slightly worse, it would be a situation it has never encountered.

And a second, independent reason: the depth camera is mounted on that moving head, so its calibration is valid for one head pose only. Move the head and the geometry is void.

Two independent constraints point at the same answer, so "look at the person while greeting and handing over, freeze while working" was not chosen — it is what survived elimination.
Good product decisions often look like this: not selected, but left standing.

10 · Method

How we know we are not fooling ourselves

This slide is not usually in a pitch. It is here because it is our real difference.

Rule one

Every claim carries its evidence level

MEASURED ran it, have the output UNVERIFIED did not
An untested inference may not be reported as a tested fact. This rule exists because we broke it.

Rule two

Report intervals, not point estimates

0/4 and 0/23 look identical in a table, but the 95 % upper bound on the true success rate differs by a factor of four. See the chart below.

Rule three

A detector counts only after someone looks

We built three cup detectors. Detection 97.6 %, jitter 6.4 mm — every number excellent. Looking at the images showed all three had locked onto the robot's own arm. Three times.

Rule four

The rule applies to other people's papers too

Oat-VLA reports 59 % vs 41 % on real robots. Recomputing from its own counts: 29/49 vs 20/49, overlapping intervals, Fisher p = 0.106 — not significant. So we cite its token reduction and not its success rate.

SAME "ZERO SUCCESSES" — 95% INTERVAL ON THE TRUE RATE 0% 50% 100% 4 trials 23 trials
Clopper-Pearson exact intervals. Four trials proves essentially nothing — the result is consistent with a true success rate as high as 60 %.
11 · Where we are

Honestly

The slides above are where we are going. This one is where we stand, including the parts that do not look good.

Working
  • Bimanual teleoperation and recording (VR and leader-arm)
  • 256 usable episodes, 116,429 frames, three cameras, 30 fps
  • Face tracking, expression, voice dialogue — as a subsystem
  • Joint-level safety layer: rejected frames 96.6 % → 0.49 %
  • Depth capture working; native recording implemented
Not working
  • Grasp success MEASURED 8/83 (95 % CI 4.3–18.1 %)
  • Locating the cup reliably — three attempts, three failures
  • Pouring, shaking, presenting — zero demonstrations recorded
  • Camera intrinsics and hand-eye calibration not done
  • The face subsystem is not connected to the arm

We put 8/83 in the pitch because that is how this project works. A robotics project that claims success has usually just not counted carefully. We counted, and we published the confidence interval too.

12 · Plan

Next, ordered by dependency

WhatWhy it is nextNeeds the robot
1Locate the cup reliably (open-vocabulary detection + human review of the images)Every later skill starts hereno
2Trial-record discipline: criteria fixed in advance, failure modes, intervalsOtherwise "improved" and "got lucky" are indistinguishableno
3Camera intrinsics + hand-eye + a saved head presetPrerequisite for any 3D reasoningyes
4Wire up the greeting stage (arm still, gaze follows)Conflicts with nothing; first demonstrable pieceyes
5Record pouring / showing / handing demonstrationsThese skills can only come from demonstrationsyes
6A safety policy for a person in the workspace (designed separately)Required before reaching toward anyoneno → yes

Item 4 deserves a separate mention: during greeting the arm is still, so it conflicts with nothing, and it is the first piece of the whole story that can be shown to someone — a robot that looks up at you, talks, and changes expression. That alone is already outside what the existing products do.

13 · Risk

The three most likely ways this fails

If this project does not work, it will probably be one of these. They are written down so they can be watched.

Risk one

Grasping never gets good enough

8/83 today. If "locate the cup reliably" cannot be solved, pouring and presenting are moot. This is the primary risk, and the reason it is item 1 on the plan.

Risk two

Demonstrations cannot teach elegance

Pour rate and stop timing live only in the demonstrations. No amount of perception tells a model how much to pour. And we have recorded none of these yet.

Risk three

Compute runs out

One SmolVLA inference takes 1329 ms and plans 50 action steps — at 30 Hz that is 1667 ms of motion, so 80 % of the budget is used, leaving 337 ms. Add a detector that must run every frame and the margin is gone.

All three have a first step that is remote, cheap, and costs no robot time: run detection offline and look at the images (one); measure the joint speed of the presenting motion before recording it (two); measure the net effect of token reduction (three). All three can be settled before the next batch of demonstrations is recorded.

14 · Close

One sentence

Existing bartender robots treat the customer as someone who collects a drink.
We are building the one that treats them as someone who is watching.

That difference decides every technical choice: why the motion is learned rather than recorded, why gaze switches by stage, why version one slides the glass instead of placing it in a hand, and why 8/83 is on a slide in this deck.

XLeRobot · 2026-09-04 · Companion documents: audience interaction plan · depth integration plan · trial-data plan
Every figure marked MEASURED is reproducible from the repository; every figure marked UNVERIFIED has not been run.
Competitor product images and the full market analysis live in 04-reports/market-analysis/ and are not published publicly.