← reports index
Assessment · 2026-09-09 · 4 days to Demo Day

Odds of placing, and what is still worth doing

Judging weights four dimensions: demo quality · technical substance · business idea · visible effort. This page places us on each of them using measured numbers only, then ranks what remains by gain ÷ cost. Everything cited has a source; where there is none it says "not measured".

If you read one thing: the seven actions that raise our odds, in order.
#ActionEffectCost
1Allow three attempts, arm_home between eachWhat the jury sees: 67% → 96.4%Near zero
2Evidence archive behind every numberProtects all four judging dimensionsDone
3Cup discrimination, detector called once before the graspTechnical substance, visible on stage1 day + on-site re-measure
4Add a target_object field; re-measure under venue lightingMakes the generalization claim checkableHalf a day on site
5Marketing evidence documentThe only entry requirement for a separate prize2 hours
6One conversation with a potential customerExplicit bonus points; turns the business case from claim into testimony15 minutes
7Knife-fight weaponNone toward a prizeFive dependent steps — do last
1, 5 and 6 together take under two days and cover both the main prize and the separate award. Details and the reasoning for each are below.

1 · Where we actually stand

Dimension What we have Where it leaks Demo quality Genuinely learned autonomy, not a script 67% unseen: 1 in 3 fails Technical Full data loop + intercontinental teleop Invisible unless narrated Business idea Bartender / service, legible use case No real customer yet Effort 82 trial batches · 130+ reports Our strongest dimension ★ 09-09 generalization test, 28 runs across 4 objects: trained yellow cup 9/10 = 90%; three unseen objects combined 8/12 = 67%
Our position on each dimension. The first three all read "something real, with a leak"; the fourth is pure strength. The bar at the bottom is the measured generalization result.

2 · Raising the odds, action 1: retry three times

The 09-09 session was 28 runs across 4 objects (object labelling by the operator who was in the room): the trained yellow cup at 9/10 = 90%; three unseen objects — blue cup 2/3, blue can 3/5, red can 3/4 — combined 8/12 = 67%.
Whichever object is on the table, a retry is the cheapest hardening available:

Attempts All fail Jury sees a success 133% 67% 210.9% 89.1% 33.6% 96.4% Table uses the unseen-object rate of 67%. With the trained cup (90% single) three attempts give 99.9%. Requirement: arm_home before every retry, or attempt 2 starts from a pose the policy has never seen and is not independent.
Same model, same data. Changing one rule — allow three attempts — takes what the jury sees from 67% to 96.4% on an unseen object, and from 90% to effectively certain on the trained cup.
Why this ranks first: it needs no training, no new data and no one on site — only three attempts written into the demo script and the narration. And it is honest: a service robot is allowed to retry, and letting the jury watch fail → auto-home → retry → succeed is more technically convincing than pretending it works first time.

3 · Raising the odds, action 2: telling cups apart (only the "detect once" version)

Open-vocabulary detection (OWLv2) has already been evaluated offline. Everything turns on keeping it outside the control loop:

WORKS — detection runs once, outside the loop Juror says "take the blue one" OWLv2 detect 1.6 s · once only pick target pose SmolVLA grasps at 30 Hz detection latency irrelevant DOES NOT WORK — detection inside the loop (B0 object tokens) detect every tick, feed into the model latency budget 337 ms OWLv2 takes 1623 ms 4.8× over budget not doable in 4 days Per-frame recall of 41.7% looks low, but over 5 frames it is 1 − 0.583⁵ ≈ 93%, and 5 frames cost 8 seconds.
The same detector is a shippable feature outside the loop and 4.8× over budget inside it. The difference is not the model — it is where you call it.
The limits that must be stated with it: the "15/15, zero errors" configuration was measured on 36 frames, one room, one lighting condition, and was never validated across scenes. Design Factory will differ in both. So this feature must be re-measured on site at T-2 before it goes on stage — a lab result cannot be demoed as-is.

4 · How the four remaining days lay out

09-10 09-11 · T-2 09-12 · T-1 09-13 · Demo remote 1 retry logic + narration 3 re-measure detector on site rehearse → freeze config evidence 2 evidence archive (done) 4 add target_object field parallel 5 marketing evidence doc  6 talk to one potential customer on site Viola captures environment full rehearsal 3 and 4 both need her on site ★ The stage is only available from 09:00 on Demo Day. T-2 and T-1 can only be run in other rooms of the same building — a different scene from the one we perform in, and our policy is scene-sensitive. Book the meeting room.
1 and 2 need nobody and are done today; 3 and 4 wait on Viola being on site; 5 and 6 run in parallel throughout. The red bar is a risk in the shape of the schedule itself: the place we can rehearse is not the place we perform.

5 · How to raise the odds — every action, ranked by gain ÷ cost

#ActionDimensionCostVerdict
1Three attempts, arm_home between eachDemo qualityNear zero, logic existsMust do — 67% → 96.4%
2Evidence archive: 4 objects × positions × rates, with keyframesTechnical + evidenceHalf a dayDone — see the evidence page
3Cup discrimination (detect once before the grasp)Technical1 day + on-site re-measureWorth it, but the on-site result decides
4Add a target_object field and re-measure under venue lightingTechnical + evidenceHalf a day on siteWorth it — object identity currently lives only in images
5Marketing evidence documentMarketing Award2 hours, need not be polishedMust do — separate prize category
6One conversation with a potential customerBonus points15 minutesMust do — cheapest point left
7Knife-fight weapon (3D-printed claws)None (side event)Design + teaching + printing + fittingLast — see below

On the knife-fight weapon

The awards document ranks this lowest and says to spend time on it only if participation barely disturbs main demo preparation. Its dependency chain is design → teach Viola to 3D-print → Space 21 arranges the teaching → print succeeds → fit to the arm — five links in four days, every one of them competing for the time of the only person on site, whose time items 3 and 4 also need.
One technical constraint too: claws are payload, and we measured today that the left shoulder-lift servo cannot hold the last 3.9° against the arm's own weight when extended. Adding mass makes that worse. If it is built, fit it to the right arm.

6 · The judgement

Top three is reachable, and it turns on whether the run on stage works.

Marketing Award is the best value on the board. All it requires is a document that need not be polished, and competitors routinely do not bother. Items 1, 5 and 6 together take under two days and cover both the main prize and the separate award.

The 90% figure holds up, and the generalization is real.
28 runs across 4 objects on 09-09 (object labelling by the operator; per-run outcomes straight from the trial records):
ObjectnSuccessFailureVoidRate
Yellow cup (in training data)1391390%
Blue cup (unseen)521267%
Blue can (unseen)632160%
Red can (unseen)431075%
Total28175677%
"90% on the same object" is exact (9/10). "80% on different cups" measures 8/12 = 67%. Quote both, and say how many of how many — the organiser warned specifically against unsupported success-rate claims, and we happen to have the support: every run has init/closure/end keyframes and a video.
The weakness is the record format, not the numbers. The trial JSON has no field for which object was used — identity exists only inside the keyframe images. The table above had to be assembled by handing 28 init frames to the operator for identification, and several wrong readings were made on the way. Add one field and this becomes a query instead of an excavation.
Sources: 05-training/trials/trials-20260909-*.json (7 batches, 17 success / 5 failure / 6 void) · 03-software/brain/components.py · owlv2-cascade-2026-09-04.html (15/15 zero errors, 1623 ms, 41.7% recall) · b0-object-tokens-2026-09-05.html (337 ms budget) · DEMO-DAY-AWARDS-AND-LOGISTICS-2026-09-10.md
Retry probabilities assume independent attempts. This page uses system fonts only and makes no external requests.