Judging weights four dimensions: demo quality · technical substance · business idea ·
visible effort. This page places us on each of them using measured numbers only, then ranks
what remains by gain ÷ cost. Everything cited has a source; where there is none it says "not measured".
If you read one thing: the seven actions that raise our odds, in order.
#
Action
Effect
Cost
1
Allow three attempts, arm_home between each
What the jury sees: 67% → 96.4%
Near zero
2
Evidence archive behind every number
Protects all four judging dimensions
Done
3
Cup discrimination, detector called once before the grasp
Technical substance, visible on stage
1 day + on-site re-measure
4
Add a target_object field; re-measure under venue lighting
Makes the generalization claim checkable
Half a day on site
5
Marketing evidence document
The only entry requirement for a separate prize
2 hours
6
One conversation with a potential customer
Explicit bonus points; turns the business case from claim into testimony
15 minutes
7
Knife-fight weapon
None toward a prize
Five dependent steps — do last
1, 5 and 6 together take under two days and cover both the main prize and the separate award.
Details and the reasoning for each are below.
1 · Where we actually stand
Our position on each dimension. The first three all read "something real, with a leak";
the fourth is pure strength. The bar at the bottom is the measured generalization result.
2 · Raising the odds, action 1: retry three times
The 09-09 session was 28 runs across 4 objects (object labelling by the operator
who was in the room): the trained yellow cup at 9/10 = 90%; three unseen objects — blue cup
2/3, blue can 3/5, red can 3/4 — combined 8/12 = 67%.
Whichever object is on the table, a retry is the cheapest hardening available:
Same model, same data. Changing one rule — allow three attempts — takes what the jury
sees from 67% to 96.4% on an unseen object, and from 90% to effectively certain on the trained cup.
Why this ranks first: it needs no training, no new data and no one on site —
only three attempts written into the demo script and the narration. And it is honest: a service
robot is allowed to retry, and letting the jury watch fail → auto-home → retry → succeed is
more technically convincing than pretending it works first time.
3 · Raising the odds, action 2: telling cups apart (only the "detect once" version)
Open-vocabulary detection (OWLv2) has already been evaluated offline. Everything turns on
keeping it outside the control loop:
The same detector is a shippable feature outside the loop and 4.8× over budget
inside it. The difference is not the model — it is where you call it.
The limits that must be stated with it: the "15/15, zero errors"
configuration was measured on 36 frames, one room, one lighting condition, and
was never validated across scenes. Design Factory will differ in both. So this feature
must be re-measured on site at T-2 before it goes on stage — a lab result cannot be demoed as-is.
4 · How the four remaining days lay out
1 and 2 need nobody and are done today; 3 and 4 wait on Viola being on site; 5 and 6 run in
parallel throughout. The red bar is a risk in the shape of the schedule itself: the place we can
rehearse is not the place we perform.
5 · How to raise the odds — every action, ranked by gain ÷ cost
#
Action
Dimension
Cost
Verdict
1
Three attempts, arm_home between each
Demo quality
Near zero, logic exists
Must do — 67% → 96.4%
2
Evidence archive: 4 objects × positions × rates, with keyframes
Technical + evidence
Half a day
Done — see the evidence page
3
Cup discrimination (detect once before the grasp)
Technical
1 day + on-site re-measure
Worth it, but the on-site result decides
4
Add a target_object field and re-measure under venue lighting
Technical + evidence
Half a day on site
Worth it — object identity currently lives only in images
5
Marketing evidence document
Marketing Award
2 hours, need not be polished
Must do — separate prize category
6
One conversation with a potential customer
Bonus points
15 minutes
Must do — cheapest point left
7
Knife-fight weapon (3D-printed claws)
None (side event)
Design + teaching + printing + fitting
Last — see below
On the knife-fight weapon
The awards document ranks this lowest and says to spend time on it only if participation
barely disturbs main demo preparation. Its dependency chain is design → teach Viola to
3D-print → Space 21 arranges the teaching → print succeeds → fit to the arm — five links in four
days, every one of them competing for the time of the only person on site, whose time items 3
and 4 also need. One technical constraint too: claws are payload, and we measured today that the left
shoulder-lift servo cannot hold the last 3.9° against the arm's own weight when extended.
Adding mass makes that worse. If it is built, fit it to the right arm.
6 · The judgement
Top three is reachable, and it turns on whether the run on stage works.
Effort is likely our strongest dimension in the room — 82 trial batches, 130+ evidenced reports, an intercontinental operation. None of that can be assembled at the last minute by anyone else
Technical substance is real (a learned policy plus a complete data loop) but will not be seen on its own — it has to be said out loud inside seven minutes
Demo quality is the dimension with the most variance — after item 1 it goes from "one in three fails" to "essentially will not fail"
Business idea turns from "we think it is useful" into "an operator says it is useful" the moment item 6 is done
Marketing Award is the best value on the board. All it requires is a document that need not be
polished, and competitors routinely do not bother. Items 1, 5 and 6 together take under two days
and cover both the main prize and the separate award.
The 90% figure holds up, and the generalization is real.
28 runs across 4 objects on 09-09 (object labelling by the operator; per-run outcomes straight from the
trial records):
Object
n
Success
Failure
Void
Rate
Yellow cup (in training data)
13
9
1
3
90%
Blue cup (unseen)
5
2
1
2
67%
Blue can (unseen)
6
3
2
1
60%
Red can (unseen)
4
3
1
0
75%
Total
28
17
5
6
77%
"90% on the same object" is exact (9/10). "80% on different cups" measures 8/12 = 67%.
Quote both, and say how many of how many — the organiser warned specifically against unsupported
success-rate claims, and we happen to have the support: every run has init/closure/end keyframes and a video.
The weakness is the record format, not the numbers. The trial JSON has no field for which
object was used — identity exists only inside the keyframe images. The table above had to be
assembled by handing 28 init frames to the operator for identification, and several wrong readings were
made on the way. Add one field and this becomes a query instead of an excavation.
Sources: 05-training/trials/trials-20260909-*.json (7 batches, 17 success / 5 failure / 6 void) ·
03-software/brain/components.py · owlv2-cascade-2026-09-04.html (15/15 zero errors, 1623 ms, 41.7% recall) ·
b0-object-tokens-2026-09-05.html (337 ms budget) · DEMO-DAY-AWARDS-AND-LOGISTICS-2026-09-10.md
Retry probabilities assume independent attempts. This page uses system fonts only and makes no external requests.