XLeRobot · 2026-09-06 · checked against the actual files on the Jetson
Copy-paste blocks, plus the two decisions you asked about.
| Repo | Episodes | Contents | On the Hub |
|---|---|---|---|
xlerobot-team/xlerobot-left-pick-cup-sep3-clean-20260905 | 54 | left arm only, one task | ✓ yes |
xlerobot-team/xlerobot-pick-cup-lr-clean-20260905 | 94 | merged & cleaned: 40 right + 54 left, two tasks | ✓ yes |
xlerobot-team/xlerobot-pick-cup-lr-sameview-20260904 | 100 | 40 right + 60 left | ✗ never uploaded |
xlerobot-team/xlerobot-pick-cup-lr-merge-20260904 | 202 | 91 right + 111 left, uncleaned | ✗ never uploaded |
xlerobot-pick-cup-lr-clean-20260905
— 94 episodes, 30,180 frames. Its two task strings are exactly:Pick up the cup with the right arm. and Pick up the cup with the left arm.
It carries no information in the left-arm episodes. Sampled from the same episode:
| at t = 3 s | at t = 8 s (mid-grasp) | |
|---|---|---|
| Left wrist | own gripper, empty table | the cup fills the frame |
| Right wrist | own gripper, empty table | identical — nothing changed |
run_policy_trials.py sends three camera streams and the
preprocessor maps them to camera1/2/3. Drop one in training and you must change the trial script
too — and a mismatch there produces no error at all, just a policy that behaves “a bit off”.python -V # needs >= 3.12 conda create -y -n lr python=3.12 && conda activate lr # only if below 3.12 pip install "lerobot[smolvla] @ git+https://github.com/huggingface/lerobot.git@22bd7a2f489b367d8df42de803b1e8c4ca63a3f9" lerobot-train --help | head -3 # no output = wrong install, stop here python -c "import lerobot; print(lerobot.__version__)" # must print 0.6.2 huggingface-cli login # paste the write token
pip install -U lerobot. The newest release on PyPI is 0.4.4, which has
no lerobot-train command at all and does not accept --dataset.eval_split.
On a machine in mainland China, run source /etc/network_turbo first.
tmux new -s sep3 # detach: Ctrl+B, release, then D · return: tmux a -t sep3
If your home connection drops, the pod keeps running and keeps billing. Without tmux the run dies with the connection and you pay for it anyway.
Option A — left arm only (54 episodes):
lerobot-train \
--policy.path=lerobot/smolvla_base \
--dataset.repo_id=xlerobot-team/xlerobot-left-pick-cup-sep3-clean-20260905 \
--rename_map='{"observation.images.head":"observation.images.camera1","observation.images.right_arm_wrist":"observation.images.camera2","observation.images.left_arm_wrist":"observation.images.camera3"}' \
--batch_size=64 --num_workers=8 --steps=20000 \
--dataset.eval_split=0 --policy.device=cuda \
--save_freq=2000 --log_freq=100 \
--output_dir=outputs/train/sep3 --job_name=sep3
Option B — merged, both arms (94 episodes):
lerobot-train \
--policy.path=lerobot/smolvla_base \
--dataset.repo_id=xlerobot-team/xlerobot-pick-cup-lr-clean-20260905 \
--rename_map='{"observation.images.head":"observation.images.camera1","observation.images.right_arm_wrist":"observation.images.camera2","observation.images.left_arm_wrist":"observation.images.camera3"}' \
--batch_size=64 --num_workers=8 --steps=20000 \
--dataset.eval_split=0 --policy.device=cuda \
--save_freq=2000 --log_freq=100 \
--output_dir=outputs/train/lr --job_name=lr
rename_map is identical for both — camera names and resolutions were checked and
match the old dataset exactly (head 424×240, both wrists 640×480). Do not edit it.
--task "Pick up the cup with the left arm." — a different wording is a task the model never
learned. (This project lost two days to that once.)
huggingface-cli upload Suyang99/xlerobot-smolvla-sep3-left \
outputs/train/sep3/checkpoints/last/pretrained_model
| Container disk | 50 GB. A SmolVLA checkpoint is ~1.8 GB and
save_freq=2000 keeps ten of them. The 10–20 GB default fills up partway through. |
| Billing | On-Demand, not Spot. A spot pod can be taken away mid-run — you lose the 20,000 steps and still pay. |
| Template | Official RunPod PyTorch 2.x. Not bare Ubuntu. |
| Pods, not Serverless | Serverless is a per-request inference endpoint, a different product entirely. |
Run the no-RTC comparison. Change exactly four things on the panel and nothing else:
| Setting | Set it to | Why |
|---|---|---|
| Async inference RTC | unticked | back to the original path — that code was never modified |
| Denoise steps | 10 (default) | the original version used the default |
| Per-trial cap | 25 s | 60 s was only needed because RTC runs at ~0.2× speed |
| Cup position | between the shoulder and the elbow | see the placement page |
Leave speed 1.0 and chunk auto alone. Re-register the batch and write the new cup
position into the initial conditions, e.g. “cup directly in front of the right shoulder, midway
between shoulder and elbow, same spot every trial”.
| Annotation button | Reported not responding while an evidence view is open. Not reproduced in a clean browser yet — the console error text (F12 → Console) would pin it down in one line, the same way the last four panel bugs were found. |
meta/info.json; upload status from meta/upload.json; task strings and per-task
episode counts from meta/tasks.parquet and the data shards; camera comparison from single frames
decoded out of the stored videos. Self-contained page — no external fonts, scripts or images.