Trainable, but not by "adding it in". lerobot 0.6.2 cannot train on multiple datasets, so the merge is done before training (both merged sets are built locally; the 202-episode one is on the Hub, sameview is pending push). On the rented GPU it is one command: ./train_smolvla_lr.sh. The recipe is identical to the deployed model b16-3cam-u20k — the only variable changed this time is the data.
About 1.3 h and ~$1 per run (RTX 4090 @ $0.75/h). Run both variants once: ~$2.5 total.
| # | Fact | Source | Consequence |
|---|---|---|---|
| 1 | Yesterday's batch is left-arm data; the deployed model is right-arm | meta/stats.json action.std: sep3 moves only dims 0–5, dims 6–11 are exactly 0; 0826 is the reverse · measured | This is not "60 more episodes of the same thing" — it is a second task. The task strings differ (…with the left arm. / …with the right arm.); SmolVLA separates them through language |
| 2 | A list in --dataset.repo_id does not work | lerobot/datasets/factory.py:128 raises NotImplementedError("The MultiLeRobotDataset isn't supported for now.") · source | The comment at configs/default.py:30-33 says a list is allowed — that comment is stale. Merge into one dataset first |
| 3 | Do not pass rename_map | The deployed checkpoint's config.json names its inputs observation.images.head / right_arm_wrist / left_arm_wrist (train_config.json has rename_map={}); the inference side run_policy_trials.py:415-417 feeds those exact names · source | The camera1/2/3 map in REMOTE-GPU-SMOLVLA-4090.md §6/§7 would give the new model input names the robot never sends |
All 13 post-gripper-swap datasets share one feature signature (17-D action/state, three camera resolutions, 30 fps, v3.0) — they merge cleanly.
| Batch | Dataset | Arm | Ep | Frames | Head view |
|---|---|---|---|---|---|
| 0826 | right-pick-cup-20260826-0042 | right | 40 | 16,487 | looking down at the table (pan 1985 / tilt 2332) |
| sep2 | right-pick-cup-sep2-* ×4 | right | 51 | 14,103 | raised ~6° (tilt 2261), sees the chairs across the table |
| sep2 | left-pick-cup-sep2-* ×2 | left | 51 | 14,293 | same |
| sep3 (yesterday) | left-pick-cup-sep3-* ×6 | left | 60 | 15,581 | looking down at the table (pan 1985 / tilt 2333) — same as 0826 |
The head-view column was checked frame by frame (one decoded head-camera frame per batch, compared side by side), not only from the numbers. The sep2 batch was shot with the head raised: the cup sits at the bottom of the frame with chairs and a wall behind it — not the deployment view. The deployment view was set by the operator to cup-grasp-v0 (preset trial = pan 2003 / tilt 2619); the datasets' recorded head state reads tilt≈2332, 287 ticks away — noted here so that, if the trials turn out view-sensitive, this is the first line to revisit.
Hence two merged datasets:
| Variant | Contents | Ep | Frames | Hub | Purpose |
|---|---|---|---|---|---|
| sameview (default) | 0826 + sep3 | 100 | 32,068 | xlerobot-team/xlerobot-pick-cup-lr-sameview-20260904 | Main run. View matches deployment, left/right balanced |
| all | all 13 | 202 | 60,464 | xlerobot-team/xlerobot-pick-cup-lr-merge-20260904 | Control. Twice the data, one mixed view |
Both task strings are in meta/tasks.parquet: index 0 = right arm, index 1 = left arm.
# On the Jetson, lerobot env. Wraps lerobot's own datasets/aggregate.py.
cd ~/Robotic_challenge/03-software/scripts
O=xlerobot-team
~/miniconda3/envs/lerobot/bin/python aggregate_datasets.py --json \
--out $O/xlerobot-pick-cup-lr-sameview-20260904 \
--repo-ids $O/xlerobot-right-pick-cup-20260826-0042 \
$O/xlerobot-left-pick-cup-sep3-20260903-{1824,1853,1906,1927,2015,2054}
A bug fixed on the way: the first attempt failed outright — with roots=None, aggregate_datasets prefers the revision-keyed cache at $HF_LEROBOT_HOME/hub/datasets--*/snapshots/<sha>/, which held only meta/ (24 KB) and no videos, so metadata validation passed 13/13 and the video copy hit FileNotFoundError. aggregate_datasets.py now passes roots and aggr_root explicitly, pointing at the complete local directories. The dashboard's "merge" button uses the same script, so it is fixed there too.
Push: LeRobotDataset(...).push_to_hub(private=<same as the reference dataset>); uplink from here ≈ 9 MB/s, so 1 GB ≈ 2 min.
Pod spec (the one already validated in REMOTE-GPU-SMOLVLA-4090.md §2): 1× RTX 4090 24 GB · ≥8 vCPU · ≥32 GB RAM · 40 GB container disk · RunPod PyTorch 2.8 / CUDA 12.8 template · ~$0.75/h.
# ① Open tmux first — a dropped home connection does not kill the run, but the pod keeps billing
tmux new -s train
# ② Environment (§3 verbatim)
nvidia-smi
pip install -U "lerobot[dataset,training,smolvla]" huggingface_hub
hf auth login # a token with read on xlerobot-team and write on Suyang99
hf auth whoami
# ③ The AV1 decode backend must be torchcodec, or data_s balloons and the GPU idles (§4)
python -c "import torch, torchcodec; print(torch.__version__, torchcodec.__version__)"
ffmpeg -decoders | grep -E "av1_cuvid"
# "undefined symbol" from torchcodec → pip install "torchcodec==0.7.0"
# ④ Put the script in place (scp from the Jetson, or paste it — it is one file)
mkdir -p /workspace && cd /workspace
# scp robomates@100.107.145.111:~/Robotic_challenge/05-training/train_smolvla_lr.sh .
chmod +x train_smolvla_lr.sh
# ⑤ 40-step smoke test (one minute). Expect step/s≈4-5, data_s<0.07, loss falling, grdn finite
BENCH=1 ./train_smolvla_lr.sh
# ⑥ The real run (default sameview; ~67-83 min)
./train_smolvla_lr.sh
# control: VARIANT=all ./train_smolvla_lr.sh
# Ctrl-B D detaches tmux; tmux attach -t train returns
# ⑦ Upload the checkpoint (create the repo, upload, confirm, only then Terminate the pod)
hf repo create xlerobot-smolvla-lr-sameview-b16-u20000 --repo-type model --private --exist-ok
hf upload Suyang99/xlerobot-smolvla-lr-sameview-b16-u20000 \
/workspace/outputs/train/smolvla_lr-sameview_b16_u20000/checkpoints/020000 checkpoints/020000 \
--repo-type model
Where every value in the script comes from: B_EFF=16 / MICRO_BATCH=16 / UPDATES=20000 / WARMUP=1000 / LR=1e-4 / AUG=1 (affine included) / eval_split=0 / seed=1000 are copied item by item from the deployed model's checkpoints/020000/pretrained_model/train_config.json. Changing any of them adds a second variable.
One thing nobody had written down: the deployed model was trained with affine augmentation on (train_config.json → tfs.affine.weight=1.0). REMOTE-GPU §10 worries that affine's ±0.05 translation could smear the cup-position cue — that worry may be valid, but the deployed model was trained with it. Kept identical this time; turning it off is the next experiment.
# On the Jetson. The panel lists models from 05-training/*/checkpoints/*/pretrained_model
# (robot_server.py:2511), so mirror the deployed model's layout. 1.3 GB; ~30 GB free on root.
cd ~/Robotic_challenge/05-training
hf download Suyang99/xlerobot-smolvla-lr-sameview-b16-u20000 \
--local-dir xlerobot-smolvla-lr-sameview-b16-u20000
Trials must name the arm explicitly. run_policy_trials.py used to take the task string from the dataset automatically — it took the first one. In the merged dataset the first one is the right arm. A guard was added today (run_policy_trials.py:_dataset_task_sentence): when a dataset has more than one task string it refuses to guess, prints both, and requires --task, verbatim:
--task "Pick up the cup with the left arm."
--task "Pick up the cup with the right arm."
The panel and the on-site console go through the task field of /api/trials/control; same rule. A wrong sentence is a different task — bartender_tasks.py:14 records the last time that went unnoticed for 31 trials over two days.
| Item | Status |
|---|---|
train_smolvla_lr.sh actually running on a 4090 | untested. No usable GPU here. bash -n and DRY=1 pass; every parameter has been used by the deployed model, only repo_id changed and rename_map removed |
| SmolVLA really separating left/right by the task sentence after merging | untested — this is the question the run answers. Risk: at the first frame, left- and right-arm episodes look almost identical; only the language differs |
| Left-arm safety floor | teleop_safety.yaml:59 has an arms.left envelope (z 0.061–0.35), but the pitch-compensated gripper length / wrist clearance were calibrated for the right arm only (the yaml comments sit under the right block). Watch the frame-reject rate before a left-arm trial |
| How much the sep2 view mismatch hurts | untested. Compare sameview vs all after both runs |
| The sameview merge itself | finished locally (100 ep / 32,068 frames, measured). The 202-ep set is on the Hub; sameview is not pushed yet — until it is, VARIANT=sameview cannot pull on the pod; use VARIANT=all or ask the remote side to push |
| Item | Time | Cost |
|---|---|---|
| Pod setup + smoke test | ~15 min | ~$0.2 |
| sameview run | 67–83 min | ~$1.0 |
| all (control) | 67–83 min | ~$1.0 |
| Upload + wrap-up | ~10 min | ~$0.1 |
| Total | ~2.5 h | ~$2.5 |
05-training/train_smolvla_lr.sh — the training script (the executable form of this page)05-training/train_smolvla_v2.sh — previous version, still defaults to the old dataset; reference only05-training/REMOTE-GPU-SMOLVLA-4090.md — pod setup, torchcodec troubleshooting, upload flow (this page reuses its §2–§4 and §11; do not use the rename_map line in its §6/§7)03-software/scripts/aggregate_datasets.py — merge (roots fixed today)03-software/scripts/run_policy_trials.py — multi-task guard added today