For: Ville (remote, currently on exchange in Taiwan) · From: the China side · 2026-09-04
You have face tracking running locally. This brief says what already exists in the repo, where your work fits, and three questions we need answered before we can plan around it.
cup_test.py instructions arrived.It is deliberately short on instructions and long on context — you know your code, we don't, and the worst outcome here is two people building the same thing twice.
You shared a 38-second screen recording (simplescreenrecorder-2026-09-03_19.04.42). Read directly off the frames:
8.4 fps detect 107.7 ms faces 1
brightness 136/255
We are not assuming anything beyond that. Everything below is either sourced from the repository or explicitly marked as a question.
Committed 2026-08-14, present on both main and china-console:
| File | Lines | What it does |
|---|---|---|
03-software/scripts/face_perception.py | 265 | Perception adapter → a single FaceObservation object. Detection + optional expression (HSEmotionONNX). No identity. |
03-software/scripts/face_follow.py | 330 | Head tracking of one primary face. Drives pan/tilt servos 7/8 directly. |
03-software/scripts/face_display.py | 168 | ESP32-S3 + OLED expressive face |
03-software/scripts/test_face_emotion.py | 113 | Camera → face/emotion smoke test |
03-software/DEMONSTRATION-INTERFACES.md | 325 | Demonstration-interface decisions |
face_follow.py's stated design goals are worth reading before you change anything — they were written against exactly the bartender use case:
vr_teleop.py or another processMeasured constants (from the source):
| Head servos | pan = 7, tilt = 8, on the left arm bus |
|---|---|
| Pan / tilt travel | 342–3754 / 1479–2617 |
| Centre / torque limit | 2048 / 200 |
| Dead zone | X 0.10, Y 0.12 (normalised image coords) |
| Primary-face selection | choose_primary(faces, previous) — sticky |
| On loss | recenter_tick() slow return; release() frees torque on exit |
Detector: OpenCV's bundled Haar cascade — no model download needed. The file header explicitly says the detector can be swapped for MediaPipe / YOLO-face without touching the head controller. That layering is good; please keep it.
There were two face detectors in this repo. The decision has now been made: YuNet.
| Where | Detector | Status |
|---|---|---|
03-software/conversational_ai/vision/face.py | YuNet (yunet.onnx, 227 KB) | ✅ Kept. Measured 18.9 ms at 640×480, 10.0 ms at 320×240 on the Jetson |
03-software/scripts/face_follow.py | OpenCV Haar cascade | ❌ Was found to be non-functional — this was not "two working options", it was one live and one dead |
Good news: you already chose YuNet too. Your cup_test.py uses models/yunet.onnx, which is the same file and the same decision we reached independently. That question is closed — we converged without needing to negotiate.
The Haar path in face_follow.py is the one that goes; it was never a working alternative.
We cannot plan around your work until we know these. Short answers are fine.
This is the one that matters most. 107.7 ms means completely different things depending on the answer:
| If it ran on… | What it implies |
|---|---|
| Your laptop | The number tells us nothing about the robot; it would have to be re-measured |
| The Jetson Orin Nano | Directly comparable to what we measured on 2026-09-04 (see the budget below) |
The robot runs on a Jetson Orin Nano Super, 7.4 GiB shared CPU/GPU memory.
The compute budget, measured on that board on 2026-09-04:
| One SmolVLA inference | 1329 ms, planning 50 action steps |
|---|---|
| 50 steps at 30 Hz | = 1667 ms of motion, so that is the budget |
| Margin | +337 ms — 80 % of the budget already used |
| All three cameras' visual cost | 177.7 ms (+59.2 ms per camera) — only 13 % of the total |
| For comparison: Grounding DINO tiny on the same board | 2657 ms — unusable online |
| Our chosen face detector, YuNet | 18.9 ms at 640×480, 10.0 ms at 320×240 |
| OWLv2, the only cup detector that works | 1623 ms — usable only as a staged "stop and look once", never per frame |
So the question for your detector is not "is it fast" in the abstract — it is whether it fits in 337 ms, or whether it can run at a lower rate than the policy. YuNet fits comfortably. A 107.7 ms detector also fits, if that number is from this board.
We can put you on the robot's Tailscale network. Tailscale is a mesh VPN, it works fine from Taiwan, and it would give you a direct route to the Jetson — no port forwarding, no tunnelling, nothing to configure on your side beyond installing the client and accepting an invite.
With that you could:
| Measure your detector on the real board | Turns Q1 from a question into a number. ⚠ Use the train310 environment: ~/miniconda3/envs/train310/bin/python. The lerobot environment's torch is CPU-only on this machine (torch.cuda.is_available() → False), so timing anything there produces a false "too slow" |
|---|---|
| Run it offline over our recorded video | 256 usable episodes / 116,429 frames of real bar-table footage already on the box (298 total, 42 of them marked trash/soak), three cameras each — head 424×240, two wrists 640×480, 30 fps. Far more representative than a webcam. ⚠ The head camera is only 424×240, which is worth knowing before you assume a detector will behave the same as on a 1080p webcam |
| Drop files straight onto the robot | http://100.107.145.111:8790/ — a drag-and-drop upload page (password-protected, same password as the dashboard) |
| See the robot's control panel | http://100.107.145.111:8770/ |
If that is useful, say so and we will send the invite. It is probably worth it even if you only use it once, because "does it run fast enough on the actual hardware" is a question that will keep coming up, and right now nobody can answer it without you.
From your message: models/efficientdet_lite0_int8.tflite for cups and models/yunet.onnx for faces, both already inside 03-software/conversational_ai/models/. Nothing to download. See §9 — we have measured that cup model and it has a problem.
YuNet is a detector, so this is detection (+ landmarks), not identity. That means the gap named in §1.2 — "recognise which customer this is" — is still open. Not a criticism: detection is what the greeting stage actually needs. It just means we should not plan around identity yet.
Meaning: can it answer "this is the same person as before" or "this is Alice" — as opposed to "there is a face at (x, y)".
Why this specific question decides everything:
face_perception.py explicitly does not do it. Identity would enable things nothing on the market does: greeting a returning customer, remembering a preference, addressing the right person when several are at the bar.We are building toward a four-stage service. Gaze appears twice — at the start and at the end — and is deliberately frozen in between:
| # | Stage | Head | Arm |
|---|---|---|---|
| 1 | Greet · converse | follows the face | still |
| 2 | Make the drink | locked | policy-driven |
| 3 | Rotate the glass to show it | follows the face | policy-driven |
| 4 | Push it toward the customer | follows the face | policy-driven |
Stage 1 is where your work lands first, and it is the easiest win in the whole project: the arm is still, so it conflicts with nothing, and "the robot looks up at you and talks to you" is already something none of the commercial bartender robots do.
Measured across all 50 recorded demonstration episodes (19,453 frames):
head_1 σ = 0.0000 range [-4.308, -4.308]
head_2 σ = 0.0000 range [49.846, 49.846]
The head never moved in any demonstration. The manipulation policy has never seen a moving head camera. Moving it during manipulation puts the input outside the distribution it was trained on. Second, independent reason: the depth camera is mounted on that moving head, so its extrinsic calibration is only valid for one head pose.
face_follow.py opens the left arm bus directly to drive servos 7/8. The manipulation policy owns that same bus while it runs, and overwrites the head every frame. They cannot run simultaneously today.
This is not a theoretical concern. From bus_guard.py's own header:
trace_teleop.py started from a terminal while a recording was[TxRxResult] Port is in use! —Our planned resolution: head control stays inside the policy process, and only the target source changes per stage (locked constant vs face-driven). One owner of the bus, no handover. If you see a better way, say so before we build it.
grep -in "human|person" policy_safety.py returns three hits, and none of them is "a person was perceived". They are: the per-joint step cap ("the largest single-tick delta any human ever commanded"), a SAFETY.md reference to a human on the stop button, and a comment about an incident.
So today, human safety = a fixed geometric envelope + someone holding the e-stop. "Push the glass toward the customer" means deliberately reaching toward a person, which is exactly what that envelope is shaped to prevent. That needs a separate design, and your perception output is the input it would need: bearing, distance, confidence, and how long a loss counts as "nobody there".
bash git checkout -b feature/face-recognition git add <your files> git commit -m "Face tracking (WIP, not wired to the arm)" git push -u origin feature/face-recognition ` ⚠ Do not commit model weights. A push failed on 2026-09-03 because of a 109 MB file — GitHub hard-rejects anything over 100 MB, and history rewriting tools are blocked in this repo. dlib's shape_predictor_68_face_landmarks.dat alone is ~100 MB. For weights: upload to HuggingFace (directly reachable from where you are), or just tell us the model name so we fetch it ourselves. If you get Tailscale access, the upload page at :8790` takes files up to 512 MB and puts them outside the repo, which is safer anyway.policy_safety.py. Its floor constraints were re-derived on 2026-09-03 (rejected-frame rate went 96.6% → 0.49%) and we are keeping it stable this week.We evaluated cup detectors offline on 36 frames sampled from three recorded datasets, then cropped and identified every single box by eye. Independently of your work.
| Detector | Boxes | Actually a cup | The target cup | Latency |
|---|---|---|---|---|
EfficientDet-Lite0 int8 — the one cup_test.py uses by default | 0/36 | — | — | 66 ms CPU |
EfficientDet-Lite2 (--large) | 7/36 | — | — | 411 ms CPU |
| OWL-ViT base-patch32 | 16/36 | 10 | 5 | 225 ms GPU |
| Grounding DINO tiny | 27/36 | 1 | 1 | 2657 ms GPU |
| OWLv2 @0.30 | 19/36 | 19 (100 %) | 16 | 1623 ms GPU |
| OWLv2 @0.30 + background difference | 15/36 | 15 | 15 — zero errors | 1623 ms GPU |
EfficientDet-Lite0-int8 produced no boxes at all on our recorded head-camera frames.
And you already knew this — the help text of your own --yolo flag says:
Two independent measurements agreeing is worth more than either alone. So: the default in cup_test.py is the model that does not work on our data. Consider making --yolo the default, or at least printing a warning.
cup_test.py opens the camera at 1280×720. But the head camera is recorded — and therefore seen by the manipulation policy — at 424×240.
That is three times the linear resolution, nine times the pixels. Small top-down cups are exactly the case where that matters most. So a detector that looks fine in the live viewer may still be useless on the frames the policy actually receives, and our 0/36 result is consistent with that.
This cuts a useful way too: the same camera clearly can deliver 1280×720. Raising the recording resolution may be much cheaper than we assumed — a configuration change rather than new hardware. <span>UNTESTED — bandwidth, disk and retraining cost all unknown</span>
cup_test.py/dev/video6, which does not exist on this machine right now. The code is fine — find_head_camera() looks the Orbbec up by name, which is the right approach. Only the comment is stale. For reference, the rest of the project addresses it by a stable path: /dev/v4l/by-path/platform-3610000.usb-usb-0:1.1:1.4-video-index0./dev/shm fallback when the camera is busy is a good design. The rest of the project has repeatedly been bitten by two programs grabbing one device — see §6.1. Publishing and reading a shared frame is exactly the right way around it.Free on this machine, no conflict with 8770 (panel), 8781 (recording) or 8790 (file drop). If you take the Tailscale route in §4.1, http://100.107.145.111:8795 would reach it directly.
Background documents, if useful: GOAL-AUDIENCE-INTERACTION.md (the audience-interaction plan, including the four-stage sequence and the safety-layer gap) · 03-software/DEMONSTRATION-INTERFACES.md (already discusses PRESENT as a distinct skill, written 2026-08-16).