XLeRobot · 2026-09-06 · checked against the actual files on the Jetson

Training run — left-arm data (0903)

Copy-paste blocks, plus the two decisions you asked about.

1 · Which dataset

RepoEpisodesContentsOn the Hub
xlerobot-team/xlerobot-left-pick-cup-sep3-clean-2026090554 left arm only, one task✓ yes
xlerobot-team/xlerobot-pick-cup-lr-clean-2026090594 merged & cleaned: 40 right + 54 left, two tasks✓ yes
xlerobot-team/xlerobot-pick-cup-lr-sameview-20260904100 40 right + 60 left✗ never uploaded
xlerobot-team/xlerobot-pick-cup-lr-merge-20260904202 91 right + 111 left, uncleaned✗ never uploaded
The merged 0903 dataset you were asking about is xlerobot-pick-cup-lr-clean-20260905 — 94 episodes, 30,180 frames. Its two task strings are exactly:
Pick up the cup with the right arm.  and  Pick up the cup with the left arm.
The other two merges are only on this Jetson — they were never pushed. Uploads happen once, at the end of recording; nothing back-fills them later. If you want one of those, upload it first.

2 · Should the right-wrist camera be dropped?

It carries no information in the left-arm episodes. Sampled from the same episode:

 at t = 3 sat t = 8 s (mid-grasp)
Left wristown gripper, empty tablethe cup fills the frame
Right wristown gripper, empty tableidentical — nothing changed
Recommendation: keep all three cameras anyway. Three reasons: Dropping it is a legitimate experiment — just not the first one, and it is not free.

3 · Install (paste as-is)

python -V                     # needs >= 3.12
conda create -y -n lr python=3.12 && conda activate lr    # only if below 3.12

pip install "lerobot[smolvla] @ git+https://github.com/huggingface/lerobot.git@22bd7a2f489b367d8df42de803b1e8c4ca63a3f9"

lerobot-train --help | head -3          # no output = wrong install, stop here
python -c "import lerobot; print(lerobot.__version__)"    # must print 0.6.2

huggingface-cli login                   # paste the write token
Do not use pip install -U lerobot. The newest release on PyPI is 0.4.4, which has no lerobot-train command at all and does not accept --dataset.eval_split.

On a machine in mainland China, run source /etc/network_turbo first.

4 · Start tmux (not optional)

tmux new -s sep3        # detach: Ctrl+B, release, then D    ·    return: tmux a -t sep3

If your home connection drops, the pod keeps running and keeps billing. Without tmux the run dies with the connection and you pay for it anyway.

5 · Train

Option A — left arm only (54 episodes):

lerobot-train \
  --policy.path=lerobot/smolvla_base \
  --dataset.repo_id=xlerobot-team/xlerobot-left-pick-cup-sep3-clean-20260905 \
  --rename_map='{"observation.images.head":"observation.images.camera1","observation.images.right_arm_wrist":"observation.images.camera2","observation.images.left_arm_wrist":"observation.images.camera3"}' \
  --batch_size=64 --num_workers=8 --steps=20000 \
  --dataset.eval_split=0 --policy.device=cuda \
  --save_freq=2000 --log_freq=100 \
  --output_dir=outputs/train/sep3 --job_name=sep3

Option B — merged, both arms (94 episodes):

lerobot-train \
  --policy.path=lerobot/smolvla_base \
  --dataset.repo_id=xlerobot-team/xlerobot-pick-cup-lr-clean-20260905 \
  --rename_map='{"observation.images.head":"observation.images.camera1","observation.images.right_arm_wrist":"observation.images.camera2","observation.images.left_arm_wrist":"observation.images.camera3"}' \
  --batch_size=64 --num_workers=8 --steps=20000 \
  --dataset.eval_split=0 --policy.device=cuda \
  --save_freq=2000 --log_freq=100 \
  --output_dir=outputs/train/lr --job_name=lr
The rename_map is identical for both — camera names and resolutions were checked and match the old dataset exactly (head 424×240, both wrists 640×480). Do not edit it.
If you train Option B, the task string matters at inference. A two-task dataset means the policy is conditioned on the sentence. When you test it you must pass the exact string, character for character: --task "Pick up the cup with the left arm." — a different wording is a task the model never learned. (This project lost two days to that once.)

6 · Push the result back

huggingface-cli upload Suyang99/xlerobot-smolvla-sep3-left \
    outputs/train/sep3/checkpoints/last/pretrained_model

7 · Pre-flight

Container disk50 GB. A SmolVLA checkpoint is ~1.8 GB and save_freq=2000 keeps ten of them. The 10–20 GB default fills up partway through.
BillingOn-Demand, not Spot. A spot pod can be taken away mid-run — you lose the 20,000 steps and still pay.
TemplateOfficial RunPod PyTorch 2.x. Not bare Ubuntu.
Pods, not ServerlessServerless is a per-request inference endpoint, a different product entirely.

8 · When you get back — the next hardware batch

Run the no-RTC comparison. Change exactly four things on the panel and nothing else:

SettingSet it toWhy
Async inference RTCuntickedback to the original path — that code was never modified
Denoise steps10 (default)the original version used the default
Per-trial cap25 s60 s was only needed because RTC runs at ~0.2× speed
Cup positionbetween the shoulder and the elbowsee the placement page

Leave speed 1.0 and chunk auto alone. Re-register the batch and write the new cup position into the initial conditions, e.g. “cup directly in front of the right shoulder, midway between shoulder and elbow, same spot every trial”.

Why this batch matters. The 2026-09-06 batch cannot answer “does RTC help?” because two things changed at once: RTC and the cup position — and the cup position turned out to be decisive. With the cup where the demonstrations put it, and RTC off, the only remaining difference from the historical batches is the cup position itself. That is a clean single-variable experiment.
Judging rule. If the arm simply ran out of time, mark it void, not failure — a time limit is not a model failure and it would pollute the denominator. If a person had to reposition the cup mid-trial, that trial is void as well, never a success.

9 · Still open

Annotation buttonReported not responding while an evidence view is open. Not reproduced in a clean browser yet — the console error text (F12 → Console) would pin it down in one line, the same way the last four panel bugs were found.
Verified on the Jetson, 2026-09-06: episode and frame counts from each dataset’s meta/info.json; upload status from meta/upload.json; task strings and per-task episode counts from meta/tasks.parquet and the data shards; camera comparison from single frames decoded out of the stored videos. Self-contained page — no external fonts, scripts or images.