← reports index
Demo Day · 2026-09-10

Demo Day: prizes, judging, schedule and roles

Sources: 00-admin/ROBOMATES-DEMO-DAY-HANDOVER-2026-09-10.md and 00-admin/DEMO-DAY-AWARDS-AND-LOGISTICS-2026-09-10.md (china-console branch). Following the repository's own discipline, confirmed and still unknown are kept apart — nothing is presented as a rule unless it is one.

Open item now closed. Both source documents carried the same caveat: the full list of prize categories was still awaiting the organiser's reply, and until it arrived nobody was to invent extra categories.
It is now confirmed: there are exactly two prize categories — the top three, and the Marketing Award. There are no others.

1 · Prizes

CategoryWhat it isStatus
Top three€7,000 total pool, split across the top 3 teamsCONFIRMED
Marketing AwardA hardware prize, under €1,000 in value. Not cash, and must not be described as "a €1,000 prize"CONFIRMED
Any other categoryNoneCONFIRMED
How the €7,000 splits between 1st/2nd/3rdNot found in repository material or organiser messagesUNKNOWN

Scores but is not a prize category

2 · The four judging dimensions

  1. How good the demo is
  2. How technically impressive the project is
  3. How good the business idea is
  4. How much effort went into it
Implication: the main-prize strategy cannot be to maximise one metric. It needs all four at once — a convincing live demonstration, real technical substance, a legible use case and customer story, and visible evidence of the work done.

3 · Demo Day schedule

09:00Design Factory stage becomes available (not earlier) 15:30Doors open / latest arrival; come earlier if setup needs longer 16:00–18:00Team presentations — originally 10 min per team (~7 min presentation and demo, ~3 min jury questions) 18:00–19:00Food while the judges deliberate 19:00–20:00Winners and the €7,000 pool announced, plus the experimental side event 20:30Event ends

The organiser mentioned that several teams have withdrawn, so there may be more stage time than planned — but do not depend on that until it is confirmed.

4 · The critical-path arithmetic

Everything below follows from one measured number: a 20,000-step training run takes about 5 hours. (ACT 04:30→08:46 = 4h16m; SmolVLA 15:04→19:58 = 4h54m, both from the training logs.) The existing left-arm model was trained on 54 episodes / 13,693 frames — 7.6 minutes of demonstration.
T-2 morning T-2 evening T-1 Demo Day record on-site data train ≈ 5 h (overnight) test + rehearse + freeze perform works data slips to T-1 morning train finishes T-1 night untested fails The 5-hour training run is what makes T-2 a hard deadline, not a preference. A model trained on T-1 data reaches the stage having never been tested in the room it has to work in.
The schedule has exactly one slack-free path. On-site data must be captured on T-2, because training consumes a whole night and the result still has to be tested before the configuration is frozen.

5 · What has to be decided, and by when

Every row is a real open question. Unanswered, each one silently removes an option later, usually at the moment there is no time left to choose.

5.1 Venue

QuestionWhy it decides somethingAskBy
How long can the DF meeting room be booked for? A few hours, a full day, or across two days?If it cannot be held overnight the robot must be re-set-up and re-calibrated each morning, and every scene-sensitive result from the day before is voidWe book it ourselvesToday
Can the robot stay assembled in the room overnight, or must it return to the DFLabs cage?Determines whether T-1 starts from a known configuration or from scratchErkkaToday
Power and network in that room?No network = no remote operation at allErkkaToday
Lighting: fixed, or can we control it? Windows, blinds, overhead lights?Our policy takes camera images as input. Lighting that changes between rehearsal and performance is a silent failure modeViola photographs it at T-2T-2
Table: same table for rehearsal and stage? Height, colour, reflectivityGrasp height comes from the demonstration data. A different table changes itErkka / ViolaT-2

5.2 On-site data — and whether any is needed at all

"How many episodes?" is the wrong question. 54 episodes is what it took to train this model from scratch. We are not doing that — we have a model that already does the task and are moving it to a different room. That is domain adaptation, and it may need zero episodes.

What the evidence actually says

FactMeasuredWhat it implies
The model already generalizes across objects8/12 = 67% on three objects that never appear in the training dataIt has not memorised the training scene. Changing what is on the table costs it something, but not everything
Colour temperature in the training data varies widelygreen/red ratio spans 0.97 – 1.33 across the 54 episodes — covering both the "neutral" and "green-tinted" reference batches in our own recordsA venue with differently-coloured lighting is probably already inside what the model has seen
Brightness does not varymean luminance 107.9, range 100–112, coefficient of variation only 2.6%★ A noticeably brighter or darker room is outside everything it has been trained on. This is the lighting axis to worry about, not colour
An on-site reading has landed outside beforethe 09-06 on-site value was 1.58, above the training maximum of 1.33The venue genuinely can fall outside the range. Measure it, do not assume
We have never fine-tuned from our own checkpointboth models were trained from lerobot/smolvla_base, 20,000 steps★ The 5-hour figure is from-scratch training time. Adapting a converged model normally needs far fewer steps — but we have no in-house measurement, so it cannot be promised

So the plan is a ladder, not a number

RungDo thisThen
0Measure the room's lighting and compare it against luminance 100–112 and g/r 0.97–1.33Inside the range → expect the model to work. Outside → expect degradation and plan for it
1Run the current model, ~10 graspsIf it holds up, record nothing and spend the day on rehearsal instead. This is the outcome to hope for
2If it degrades: record 10–15 episodes and fine-tune from our own checkpoint, not from the base model~20 minutes of recording. This is the step we have never done, so time the first run and write the number down
3Only if rung 2 fails: record up to 40 and retrain properlyCosts the whole night. This is the branch the schedule cannot absorb twice
What to decide today, before anyone is on site: not "how many episodes", but who runs rung 0 and rung 1, and what result sends us to rung 2. Pick the threshold in advance — "fewer than 6 of 10 succeed" — so the decision on the day is a reading, not an argument.
Remaining questionsWhat we knowDecide
Who resets the scene between episodes?Viola. The organisers said they cannot stay for thisConfirm she is free for a continuous 20–60 minutes, not fragments
Which object do we record?Yellow cup = 9/10 on the current model. Unseen objects = 8/12Record the object that goes on stage. Recording one and demoing another discards the advantage
Is the cloud GPU reserved?The Jetson cannot train — its torch has no CUDAReserve it before T-2, not after the data exists

5.3 Training and testing

QuestionWhy it mattersDecide
Where does training run, and is it available that night?The Jetson cannot train — its torch has no CUDA. Training happens on a rented cloud GPUConfirm the instance is reserved before T-2, not after the data exists
How much time is reserved for testing after training?A model that has never been run on hardware is not a demo, it is a hopeReserve at least 2 hours on T-1 for hardware trials, and treat a bad result as a reason to fall back to the current model
What is the fallback if the new model is worse?Fine-tuning on a small on-site set can degrade a working policyKeep the current checkpoint loadable and decide by a stated time on T-1, not by feel on the day
When exactly is the configuration frozen?"After the rehearsal" is not a timeName an hour on T-1. After it, no model, pose or camera changes

5.4 Features — decide whether to build at all

QuestionStateDecide
Cup discrimination — do we do it?Feasible only outside the control loop (detect once before the grasp; inside it, 1623 ms against a 337 ms budget). The "15/15, zero errors" result was measured on 36 frames, one room, one lighting condition and was never validated across scenesDecide today whether it is in scope. If yes it must be re-measured on site at T-2 before it is allowed on stage — a lab number cannot be demoed
Three-attempt retryNeeds no training and no one on site. Takes what the jury sees from 67% to 96.4% on an unseen objectDo it. There is no argument against it
Knife-fight weapon (3D-printed claws)Five dependent steps in the remaining days, all competing for the only person on site. And claws are payload on an arm whose shoulder-lift already cannot hold the last 3.9° when extendedOnly if it costs nothing elsewhere. If built, fit it to the right arm

5.5 Hardware

QuestionWhy it mattersAsk
Will someone inspect the left shoulder-lift servo, and when?It stalls 3.9° short of an extended pose at only 40 °C. We need the answer before freezing the configuration, so we know whether to design the demo around a weaker jointErkka — the note is drafted and ready to send
Is there a spare Feetech STS/SCS servo anywhere nearby?Not to fit now — so that a failure in the last two days is not a blocker with only one person on siteErkka
Who does the hands-on left-vs-right resistance comparison?It is the single check that separates "the joint is below spec" from "the pose is beyond this servo size". Torque is already disabled — Viola can do it in two minutesViola, any time

5.6 The performance itself

QuestionDecide
Which object goes on the table in front of the jury?The trained object is the safe choice at 9/10. An unseen object is the more impressive claim at 8/12. Pick one and rehearse only that
How many attempts do we allow ourselves on stage, and who says so out loud?Announce it before the run — "it gets three tries, like a bartender would" — so a retry reads as designed rather than as a save
Who narrates while the arm moves?The 7 minutes are the only chance to make the invisible technical substance visible
Live or recorded, and who decides?Live scores far higher. Name the person and the moment at which the fallback video is played instead

6 · The T-2 sequence, in order, with the dependencies

This is the part that has been under-planned. Recording data is the last step, not the first — several things have to be right before a single episode is worth keeping.

#StepTimeWhy it blocks what follows
1Organisers deliver, power and network the robotunknownNothing starts until this is done. Get an arrival time from Erkka — the whole day's arithmetic hangs off it
2Physical setup: arms, cameras, cabling~30 min—
3Re-check the USB camera paths10 min★ The camera paths are resolved by USB topology. Re-plugging into different ports changes them and the code stops finding the cameras — this has happened before and took a whole session to find. Verify all four streams appear before anything else
4Table measurement and alignment20–30 minGrasp height comes from the demonstration data. A table at a different height or position invalidates every stored pose. Measure before deciding whether to re-record
5Calibration — only if step 4 says it is needed~10 min/armSee the warning below. Do not recalibrate reflexively
6Camera geometry photographed and documented; lighting recorded15 minNeeded by the remote team to judge whether the scene matches training
7Sanity run: current model, a handful of grasps20 minTells you whether new data is needed at all. If the existing model still works in the room, the rest of the day is free
8Re-measure the cup detector if discrimination is in scope30 minThe zero-error result was one room, one lighting condition. Unvalidated elsewhere
9Record episodes45–60 min~12 s per episode plus a scene reset; 40–60 episodes. Needs Viola for a continuous block
10Upload, verify, start training15 min + 5 hMust start early enough to finish overnight

Steps 1–8 are roughly 2–2.5 hours before recording even begins. Add an hour of things going wrong. If the robot arrives in the afternoon, training does not start that night — and that is the scenario the critical path above says we cannot afford.

★ Recalibration has a cost that is easy to miss ★
Joint angles in the recorded data and in every stored pose are relative to the current calibration. Recalibrating shifts that frame, so afterwards: So: measure first (step 4), and only calibrate if the measurement says the arm cannot reach the poses it needs. If calibration does happen, budget the re-recording that follows it — it is not 10 minutes, it is 10 minutes plus an hour.

7 · T-1 and Demo Day

T-1

StepTimeNote
Restore yesterday's configuration30 minTrivial if the room was held overnight; a rebuild if not — see §5.1
Hardware trials of the newly trained model≥2 hA model that has never run on hardware is not a demo
Decision point: new model or the current one—Name the hour in advance. Fine-tuning on a small on-site set can make a working policy worse
Full rehearsal, timed1 hIncluding the narration, not just the arm
Teleop fallback test · backup video test30 minBoth fallbacks must be known-working before the freeze
Freeze the configuration—After this: no model, pose, camera or lighting changes

Demo Day

8 · Who does what

PersonWhereResponsibilities
ViolaOn site in Finland
(the only member present)
Scene setup and reset · cup and prop placement · onsite calibration and data collection under remote instruction · photographing and measuring lighting, background, table and camera geometry · simple physical adjustments · physical E-stop and safety · liaison with the organisers
Suyang and VilleRemote — China and TaiwanTechnical work and operation of the robot
Viola is not responsible for transporting the robot, nor for debugging the system independently. The remote team must reduce whatever it needs into a very short checklist of physical actions that requires no understanding of the software stack.

9 · Venue and organiser support

The organisers can: bring the robot to the venue · plug it in · provide internet.
They most likely cannot: stay to reset the scene repeatedly, or supervise data collection.

10 · Evidence rules that affect scoring

Do not make unsupported quantitative or generalization claims.
Anything of the form "90% success", "different cups", "different positions" or "language-conditioned selection" must have physical evidence behind it. Strong evidence should make clear that the autonomy is neither replay nor teleoperation, and should genuinely vary the target object or its position wherever generalization is claimed.

The organiser stated that a live demo scores far higher than a video. Prefer the live run, and keep a best-performance recording as fallback.

11 · Hardware risk

The left shoulder-lift servo has repeatedly overheated under sustained or held load. It still works. A replacement servo was already fitted at this joint once, so another swap is not assumed to be the answer. Erkka said he would try to find time or find someone to inspect it before Demo Day, with no guarantee.
Mitigation: keep the thermal guard and idle auto-park active, and avoid sustained high-load held poses. This is not Viola's repair responsibility.
New measurement, 2026-09-09: the same servo stalls at only 40 °C. Commanded to the pour pose it moved 43 ticks and then stopped, still 45 ticks (3.9°) short, and did not recover — while the same servo reaches the home pose with 0.0° error. So this is not thermal derating; it looks load-dependent.
Direct conflict with the demo: the full pour sequence requires the left arm to hold an extended pose, which is exactly the duty cycle the mitigation above tells us to avoid. A separate note has been sent to the organisers.

12 · Demo strategy

Target: top three plus the Marketing Award.

Do not pitch this as "a robot that picks up a cup".
The story is: remote human operation → demonstration data → policy training → increasing autonomy → human takeover where autonomy is insufficient.

What the Marketing Award needs

A standalone marketing evidence document, which the organiser explicitly said need not be polished: where the project was marketed · how many posts or actions · views, impressions or reach where available · screenshots and links.

13 · Source-of-truth precedence

  1. Logistics and who is present follow the handover document
  2. DEMO-DAY-AWARDS-AND-LOGISTICS-2026-09-10.md is the full plan
  3. Read the latest CURRENT-STATUS.md before quoting model or autonomy performance
  4. Do not reuse the Final Check-up line saying there are zero team members on site — it is superseded; Viola is there
  5. Do not invent unconfirmed prize categories (now confirmed: top three and Marketing Award only)
Sources: 00-admin/ROBOMATES-DEMO-DAY-HANDOVER-2026-09-10.md · 00-admin/DEMO-DAY-AWARDS-AND-LOGISTICS-2026-09-10.md (china-console branch)
Prize-category scope confirmed by the operator on 2026-09-09. This page uses system fonts only and makes no external requests.