← reports index
Training · 2026-09-04

Training on yesterday's batch: merge, rent a GPU, run once

2026-09-04 · Evidence levels: measured = run on this machine; source = file:line; untested = not run, do not treat as a result.

Verdict

Trainable, but not by "adding it in". lerobot 0.6.2 cannot train on multiple datasets, so the merge is done before training (both merged sets are built locally; the 202-episode one is on the Hub, sameview is pending push). On the rented GPU it is one command: ./train_smolvla_lr.sh. The recipe is identical to the deployed model b16-3cam-u20k — the only variable changed this time is the data.

About 1.3 h and ~$1 per run (RTX 4090 @ $0.75/h). Run both variants once: ~$2.5 total.

0. Three things that would waste the run (read first)

#FactSourceConsequence
1Yesterday's batch is left-arm data; the deployed model is right-armmeta/stats.json action.std: sep3 moves only dims 0–5, dims 6–11 are exactly 0; 0826 is the reverse · measuredThis is not "60 more episodes of the same thing" — it is a second task. The task strings differ (…with the left arm. / …with the right arm.); SmolVLA separates them through language
2A list in --dataset.repo_id does not worklerobot/datasets/factory.py:128 raises NotImplementedError("The MultiLeRobotDataset isn't supported for now.") · sourceThe comment at configs/default.py:30-33 says a list is allowed — that comment is stale. Merge into one dataset first
3Do not pass rename_mapThe deployed checkpoint's config.json names its inputs observation.images.head / right_arm_wrist / left_arm_wrist (train_config.json has rename_map={}); the inference side run_policy_trials.py:415-417 feeds those exact names · sourceThe camera1/2/3 map in REMOTE-GPU-SMOLVLA-4090.md §6/§7 would give the new model input names the robot never sends

1. Data inventory (measured)

All 13 post-gripper-swap datasets share one feature signature (17-D action/state, three camera resolutions, 30 fps, v3.0) — they merge cleanly.

BatchDatasetArmEpFramesHead view
0826right-pick-cup-20260826-0042right4016,487looking down at the table (pan 1985 / tilt 2332)
sep2right-pick-cup-sep2-* ×4right5114,103raised ~6° (tilt 2261), sees the chairs across the table
sep2left-pick-cup-sep2-* ×2left5114,293same
sep3 (yesterday)left-pick-cup-sep3-* ×6left6015,581looking down at the table (pan 1985 / tilt 2333) — same as 0826

The head-view column was checked frame by frame (one decoded head-camera frame per batch, compared side by side), not only from the numbers. The sep2 batch was shot with the head raised: the cup sits at the bottom of the frame with chairs and a wall behind it — not the deployment view. The deployment view was set by the operator to cup-grasp-v0 (preset trial = pan 2003 / tilt 2619); the datasets' recorded head state reads tilt≈2332, 287 ticks away — noted here so that, if the trials turn out view-sensitive, this is the first line to revisit.

Hence two merged datasets:

VariantContentsEpFramesHubPurpose
sameview (default)0826 + sep310032,068xlerobot-team/xlerobot-pick-cup-lr-sameview-20260904Main run. View matches deployment, left/right balanced
allall 1320260,464xlerobot-team/xlerobot-pick-cup-lr-merge-20260904Control. Twice the data, one mixed view

Both task strings are in meta/tasks.parquet: index 0 = right arm, index 1 = left arm.

2. How the merge was made (measured on this machine)

# On the Jetson, lerobot env. Wraps lerobot's own datasets/aggregate.py.
cd ~/Robotic_challenge/03-software/scripts
O=xlerobot-team
~/miniconda3/envs/lerobot/bin/python aggregate_datasets.py --json \
  --out $O/xlerobot-pick-cup-lr-sameview-20260904 \
  --repo-ids $O/xlerobot-right-pick-cup-20260826-0042 \
             $O/xlerobot-left-pick-cup-sep3-20260903-{1824,1853,1906,1927,2015,2054}

A bug fixed on the way: the first attempt failed outright — with roots=None, aggregate_datasets prefers the revision-keyed cache at $HF_LEROBOT_HOME/hub/datasets--*/snapshots/<sha>/, which held only meta/ (24 KB) and no videos, so metadata validation passed 13/13 and the video copy hit FileNotFoundError. aggregate_datasets.py now passes roots and aggr_root explicitly, pointing at the complete local directories. The dashboard's "merge" button uses the same script, so it is fixed there too.

Push: LeRobotDataset(...).push_to_hub(private=<same as the reference dataset>); uplink from here ≈ 9 MB/s, so 1 GB ≈ 2 min.

3. On the rented GPU, step by step

Pod spec (the one already validated in REMOTE-GPU-SMOLVLA-4090.md §2): 1× RTX 4090 24 GB · ≥8 vCPU · ≥32 GB RAM · 40 GB container disk · RunPod PyTorch 2.8 / CUDA 12.8 template · ~$0.75/h.

# ① Open tmux first — a dropped home connection does not kill the run, but the pod keeps billing
tmux new -s train

# ② Environment (§3 verbatim)
nvidia-smi
pip install -U "lerobot[dataset,training,smolvla]" huggingface_hub
hf auth login          # a token with read on xlerobot-team and write on Suyang99
hf auth whoami

# ③ The AV1 decode backend must be torchcodec, or data_s balloons and the GPU idles (§4)
python -c "import torch, torchcodec; print(torch.__version__, torchcodec.__version__)"
ffmpeg -decoders | grep -E "av1_cuvid"
# "undefined symbol" from torchcodec → pip install "torchcodec==0.7.0"

# ④ Put the script in place (scp from the Jetson, or paste it — it is one file)
mkdir -p /workspace && cd /workspace
# scp robomates@100.107.145.111:~/Robotic_challenge/05-training/train_smolvla_lr.sh .
chmod +x train_smolvla_lr.sh

# ⑤ 40-step smoke test (one minute). Expect step/s≈4-5, data_s<0.07, loss falling, grdn finite
BENCH=1 ./train_smolvla_lr.sh

# ⑥ The real run (default sameview; ~67-83 min)
./train_smolvla_lr.sh
#    control:  VARIANT=all ./train_smolvla_lr.sh
#    Ctrl-B D detaches tmux; tmux attach -t train returns

# ⑦ Upload the checkpoint (create the repo, upload, confirm, only then Terminate the pod)
hf repo create xlerobot-smolvla-lr-sameview-b16-u20000 --repo-type model --private --exist-ok
hf upload Suyang99/xlerobot-smolvla-lr-sameview-b16-u20000 \
   /workspace/outputs/train/smolvla_lr-sameview_b16_u20000/checkpoints/020000 checkpoints/020000 \
   --repo-type model

Where every value in the script comes from: B_EFF=16 / MICRO_BATCH=16 / UPDATES=20000 / WARMUP=1000 / LR=1e-4 / AUG=1 (affine included) / eval_split=0 / seed=1000 are copied item by item from the deployed model's checkpoints/020000/pretrained_model/train_config.json. Changing any of them adds a second variable.

One thing nobody had written down: the deployed model was trained with affine augmentation on (train_config.json → tfs.affine.weight=1.0). REMOTE-GPU §10 worries that affine's ±0.05 translation could smear the cup-position cue — that worry may be valid, but the deployed model was trained with it. Kept identical this time; turning it off is the next experiment.

4. Pull to the Jetson and run trials

# On the Jetson. The panel lists models from 05-training/*/checkpoints/*/pretrained_model
# (robot_server.py:2511), so mirror the deployed model's layout. 1.3 GB; ~30 GB free on root.
cd ~/Robotic_challenge/05-training
hf download Suyang99/xlerobot-smolvla-lr-sameview-b16-u20000 \
   --local-dir xlerobot-smolvla-lr-sameview-b16-u20000

Trials must name the arm explicitly. run_policy_trials.py used to take the task string from the dataset automatically — it took the first one. In the merged dataset the first one is the right arm. A guard was added today (run_policy_trials.py:_dataset_task_sentence): when a dataset has more than one task string it refuses to guess, prints both, and requires --task, verbatim:

--task "Pick up the cup with the left arm."
--task "Pick up the cup with the right arm."

The panel and the on-site console go through the task field of /api/trials/control; same rule. A wrong sentence is a different task — bartender_tasks.py:14 records the last time that went unnoticed for 31 trials over two days.

5. What is not verified (decide how much to trust after reading this)

ItemStatus
train_smolvla_lr.sh actually running on a 4090untested. No usable GPU here. bash -n and DRY=1 pass; every parameter has been used by the deployed model, only repo_id changed and rename_map removed
SmolVLA really separating left/right by the task sentence after merginguntested — this is the question the run answers. Risk: at the first frame, left- and right-arm episodes look almost identical; only the language differs
Left-arm safety floorteleop_safety.yaml:59 has an arms.left envelope (z 0.061–0.35), but the pitch-compensated gripper length / wrist clearance were calibrated for the right arm only (the yaml comments sit under the right block). Watch the frame-reject rate before a left-arm trial
How much the sep2 view mismatch hurtsuntested. Compare sameview vs all after both runs
The sameview merge itselffinished locally (100 ep / 32,068 frames, measured). The 202-ep set is on the Hub; sameview is not pushed yet — until it is, VARIANT=sameview cannot pull on the pod; use VARIANT=all or ask the remote side to push

6. Cost

ItemTimeCost
Pod setup + smoke test~15 min~$0.2
sameview run67–83 min~$1.0
all (control)67–83 min~$1.0
Upload + wrap-up~10 min~$0.1
Total~2.5 h~$2.5

Related files