← reports index
2026-09-04 · brief for Ville · updated after his cup_test.py

Brief — Face recognition & audience interaction

For: Ville (remote, currently on exchange in Taiwan) · From: the China side · 2026-09-04

Practical notes given where you are: **GitHub and HuggingFace are directly reachable from
Taiwan**, no tunnelling needed — so pushing a branch or uploading weights is straightforward.
And see §4.1: we can give you direct network access to the robot, which turns the biggest
open question below into something you can just measure yourself.

You have face tracking running locally. This brief says what already exists in the repo, where your work fits, and three questions we need answered before we can plan around it.

Updated 2026-09-04 after your cup_test.py instructions arrived.
Q2 and Q3 are now answered — see §4. Q1 (which machine the 107.7 ms was measured on)
is still open, and §4.1 offers a way to settle it by measurement rather than memory.
§9 is new and is the part most worth your time: we ran an independent evaluation of cup
detectors on our recorded frames, and it agrees with a note you left in your own code.

It is deliberately short on instructions and long on context — you know your code, we don't, and the worst outcome here is two people building the same thing twice.


1. What we saw

You shared a 38-second screen recording (simplescreenrecorder-2026-09-03_19.04.42). Read directly off the frames:

8.4 fps    detect 107.7 ms    faces 1
brightness 136/255

We are not assuming anything beyond that. Everything below is either sourced from the repository or explicitly marked as a question.


2. What is already in the repo (so you don't rebuild it)

Committed 2026-08-14, present on both main and china-console:

FileLinesWhat it does
03-software/scripts/face_perception.py265Perception adapter → a single FaceObservation object. Detection + optional expression (HSEmotionONNX). No identity.
03-software/scripts/face_follow.py330Head tracking of one primary face. Drives pan/tilt servos 7/8 directly.
03-software/scripts/face_display.py168ESP32-S3 + OLED expressive face
03-software/scripts/test_face_emotion.py113Camera → face/emotion smoke test
03-software/DEMONSTRATION-INTERFACES.md325Demonstration-interface decisions

face_follow.py's stated design goals are worth reading before you change anything — they were written against exactly the bartender use case:

- follow ONE primary face instead of jumping between people
- proportional motion with a dead zone, smoothing and hard per-tick limits
- no servo chatter when a face is already near image centre
- after losing the person, wait, then gently return to centre
- never share a serial bus with vr_teleop.py or another process

Measured constants (from the source):

Head servospan = 7, tilt = 8, on the left arm bus
Pan / tilt travel342–3754 / 1479–2617
Centre / torque limit2048 / 200
Dead zoneX 0.10, Y 0.12 (normalised image coords)
Primary-face selectionchoose_primary(faces, previous) — sticky
On lossrecenter_tick() slow return; release() frees torque on exit

Detector: OpenCV's bundled Haar cascade — no model download needed. The file header explicitly says the detector can be swapped for MediaPipe / YOLO-face without touching the head controller. That layering is good; please keep it.


3. ⚠ Detector convergence — decided 2026-09-04

There were two face detectors in this repo. The decision has now been made: YuNet.

WhereDetectorStatus
03-software/conversational_ai/vision/face.pyYuNet (yunet.onnx, 227 KB)✅ Kept. Measured 18.9 ms at 640×480, 10.0 ms at 320×240 on the Jetson
03-software/scripts/face_follow.pyOpenCV Haar cascade❌ Was found to be non-functional — this was not "two working options", it was one live and one dead

Good news: you already chose YuNet too. Your cup_test.py uses models/yunet.onnx, which is the same file and the same decision we reached independently. That question is closed — we converged without needing to negotiate.

The Haar path in face_follow.py is the one that goes; it was never a working alternative.


4. The three questions we need answered

We cannot plan around your work until we know these. Short answers are fine.

Q1 · What machine was that recording made on?

This is the one that matters most. 107.7 ms means completely different things depending on the answer:

If it ran on…What it implies
Your laptopThe number tells us nothing about the robot; it would have to be re-measured
The Jetson Orin NanoDirectly comparable to what we measured on 2026-09-04 (see the budget below)

The robot runs on a Jetson Orin Nano Super, 7.4 GiB shared CPU/GPU memory.

The compute budget, measured on that board on 2026-09-04:

One SmolVLA inference1329 ms, planning 50 action steps
50 steps at 30 Hz= 1667 ms of motion, so that is the budget
Margin+337 ms — 80 % of the budget already used
All three cameras' visual cost177.7 ms (+59.2 ms per camera) — only 13 % of the total
For comparison: Grounding DINO tiny on the same board2657 ms — unusable online
Our chosen face detector, YuNet18.9 ms at 640×480, 10.0 ms at 320×240
OWLv2, the only cup detector that works1623 ms — usable only as a staged "stop and look once", never per frame

So the question for your detector is not "is it fast" in the abstract — it is whether it fits in 337 ms, or whether it can run at a lower rate than the policy. YuNet fits comfortably. A 107.7 ms detector also fits, if that number is from this board.

4.1 ★ Better than answering Q1: measure it on the actual board ★

We can put you on the robot's Tailscale network. Tailscale is a mesh VPN, it works fine from Taiwan, and it would give you a direct route to the Jetson — no port forwarding, no tunnelling, nothing to configure on your side beyond installing the client and accepting an invite.

With that you could:

Measure your detector on the real boardTurns Q1 from a question into a number. ⚠ Use the train310 environment: ~/miniconda3/envs/train310/bin/python. The lerobot environment's torch is CPU-only on this machine (torch.cuda.is_available() → False), so timing anything there produces a false "too slow"
Run it offline over our recorded video256 usable episodes / 116,429 frames of real bar-table footage already on the box (298 total, 42 of them marked trash/soak), three cameras each — head 424×240, two wrists 640×480, 30 fps. Far more representative than a webcam. ⚠ The head camera is only 424×240, which is worth knowing before you assume a detector will behave the same as on a 1080p webcam
Drop files straight onto the robothttp://100.107.145.111:8790/ — a drag-and-drop upload page (password-protected, same password as the dashboard)
See the robot's control panelhttp://100.107.145.111:8770/

If that is useful, say so and we will send the invite. It is probably worth it even if you only use it once, because "does it run fast enough on the actual hardware" is a question that will keep coming up, and right now nobody can answer it without you.

Q2 · Which detector / model? — ✅ answered

From your message: models/efficientdet_lite0_int8.tflite for cups and models/yunet.onnx for faces, both already inside 03-software/conversational_ai/models/. Nothing to download. See §9 — we have measured that cup model and it has a problem.

Q3 · Does it do identity, or detection + landmarks? — ✅ answered

YuNet is a detector, so this is detection (+ landmarks), not identity. That means the gap named in §1.2 — "recognise which customer this is" — is still open. Not a criticism: detection is what the greeting stage actually needs. It just means we should not plan around identity yet.

Meaning: can it answer "this is the same person as before" or "this is Alice" — as opposed to "there is a face at (x, y)".

Why this specific question decides everything:


5. Where this fits in the product

We are building toward a four-stage service. Gaze appears twice — at the start and at the end — and is deliberately frozen in between:

#StageHeadArm
1Greet · conversefollows the facestill
2Make the drinklockedpolicy-driven
3Rotate the glass to show itfollows the facepolicy-driven
4Push it toward the customerfollows the facepolicy-driven

Stage 1 is where your work lands first, and it is the easiest win in the whole project: the arm is still, so it conflicts with nothing, and "the robot looks up at you and talks to you" is already something none of the commercial bartender robots do.

Why the head must be frozen during stage 2 — this is not a preference

Measured across all 50 recorded demonstration episodes (19,453 frames):

head_1    σ = 0.0000    range [-4.308, -4.308]
head_2    σ = 0.0000    range [49.846, 49.846]

The head never moved in any demonstration. The manipulation policy has never seen a moving head camera. Moving it during manipulation puts the input outside the distribution it was trained on. Second, independent reason: the depth camera is mounted on that moving head, so its extrinsic calibration is only valid for one head pose.


6. Two constraints you will hit

6.1 The serial bus is exclusive

face_follow.py opens the left arm bus directly to drive servos 7/8. The manipulation policy owns that same bus while it runs, and overwrites the head every frame. They cannot run simultaneously today.

This is not a theoretical concern. From bus_guard.py's own header:

on 2026-08-17 at 23:16 a trace_teleop.py started from a terminal while a recording was
running, and the recording died with [TxRxResult] Port is in use! —
losing the episode and dropping both arms onto the table.

Our planned resolution: head control stays inside the policy process, and only the target source changes per stage (locked constant vs face-driven). One owner of the bus, no handover. If you see a better way, say so before we build it.

6.2 The safety layer has no concept of a person

grep -in "human|person" policy_safety.py returns three hits, and none of them is "a person was perceived". They are: the per-joint step cap ("the largest single-tick delta any human ever commanded"), a SAFETY.md reference to a human on the stop button, and a comment about an incident.

So today, human safety = a fixed geometric envelope + someone holding the e-stop. "Push the glass toward the customer" means deliberately reaching toward a person, which is exactly what that envelope is shaped to prevent. That needs a separate design, and your perception output is the input it would need: bearing, distance, confidence, and how long a loss counts as "nobody there".


7. What would help us most, in order

  1. Answer Q1–Q3 (above). Two lines each is enough — or take the Tailscale route in §4.1 and measure Q1 directly, which is strictly better.
  2. Push the code to a branch so we can read it: ``bash git checkout -b feature/face-recognition git add <your files> git commit -m "Face tracking (WIP, not wired to the arm)" git push -u origin feature/face-recognition ` ⚠ Do not commit model weights. A push failed on 2026-09-03 because of a 109 MB file — GitHub hard-rejects anything over 100 MB, and history rewriting tools are blocked in this repo. dlib's shape_predictor_68_face_landmarks.dat alone is ~100 MB. For weights: upload to HuggingFace (directly reachable from where you are), or just tell us the model name so we fetch it ourselves. If you get Tailscale access, the upload page at :8790` takes files up to 512 MB and puts them outside the repo, which is safer anyway.
  3. If you have an opinion on §3 (which of the three detectors should survive), tell us. You are the only one who has run one of them on real video.

8. What we are NOT asking you to do


9. ★ What we measured on the cup detector — this is the part to read ★

We evaluated cup detectors offline on 36 frames sampled from three recorded datasets, then cropped and identified every single box by eye. Independently of your work.

DetectorBoxesActually a cupThe target cupLatency
EfficientDet-Lite0 int8 — the one cup_test.py uses by default0/36——66 ms CPU
EfficientDet-Lite2 (--large)7/36——411 ms CPU
OWL-ViT base-patch3216/36105225 ms GPU
Grounding DINO tiny27/36112657 ms GPU
OWLv2 @0.3019/3619 (100 %)161623 ms GPU
OWLv2 @0.30 + background difference15/361515 — zero errors1623 ms GPU

EfficientDet-Lite0-int8 produced no boxes at all on our recorded head-camera frames.

And you already knew this — the help text of your own --yolo flag says:

"Much better on small/top-down cups: 0.80 vs nothing on the same head frame."

Two independent measurements agreeing is worth more than either alone. So: the default in cup_test.py is the model that does not work on our data. Consider making --yolo the default, or at least printing a warning.

9.1 One thing that may explain why it looks better live than it measures

cup_test.py opens the camera at 1280×720. But the head camera is recorded — and therefore seen by the manipulation policy — at 424×240.

That is three times the linear resolution, nine times the pixels. Small top-down cups are exactly the case where that matters most. So a detector that looks fine in the live viewer may still be useless on the frames the policy actually receives, and our 0/36 result is consistent with that.

This cuts a useful way too: the same camera clearly can deliver 1280×720. Raising the recording resolution may be much cheaper than we assumed — a configuration change rather than new hardware. <span>UNTESTED — bandwidth, disk and retraining cost all unknown</span>

9.2 Two small things in cup_test.py

  1. The docstring says /dev/video6, which does not exist on this machine right now. The code is fine — find_head_camera() looks the Orbbec up by name, which is the right approach. Only the comment is stale. For reference, the rest of the project addresses it by a stable path: /dev/v4l/by-path/platform-3610000.usb-usb-0:1.1:1.4-video-index0.
  2. The /dev/shm fallback when the camera is busy is a good design. The rest of the project has repeatedly been bitten by two programs grabbing one device — see §6.1. Publishing and reading a shared frame is exactly the right way around it.

9.3 Port 8795

Free on this machine, no conflict with 8770 (panel), 8781 (recording) or 8790 (file drop). If you take the Tailscale route in §4.1, http://100.107.145.111:8795 would reach it directly.


Background documents, if useful: GOAL-AUDIENCE-INTERACTION.md (the audience-interaction plan, including the four-stage sequence and the safety-layer gap) · 03-software/DEMONSTRATION-INTERFACES.md (already discusses PRESENT as a distinct skill, written 2026-08-16).