Skip to content

Train a reference policy

Current open-core protocol

The open core trains ACT, diffusion policy (DP) and chunked behavioral cloning (BC). Collect and audit data with the source campaign, then train an individual model from its generated task dataset:

WASMAN_ASSET_PROFILE=open-procedural-v1 .venv/bin/python scripts/train_revision_policy.py \
  --task PressButton --model ACT --seed 17 \
  --data artifacts/open-core/data/PressButton \
  --output-dir artifacts/retrained-button-act17

For this individual command, retain the raw task RGB until training finishes; the complete campaign archives it losslessly after fitting its reference models. Use restore_revision_rgb.py when working from those local lossless archives.

The same entry point accepts --model DP or --model BC. Use a fresh output directory for each run. The frozen primary protocol requires seeds 17, 43 and 101, the identical 76/4 whole-episode split and 320,000 sampled windows, with the final-budget checkpoint for every family. The trainer checks the open asset profile, runtime hashes and physical dataset audit; it rejects pilot data for primary training. See the protocol for full cohorts and the reproduction record for completed checks and cross-workstation limits. A task package's final checkpoints do not contain the optimizer state needed to resume their historical intermediate updates.

Historical CAD-profile training

Training is optional when evaluating an existing compatible checkpoint. The source release provides trainers and configuration snapshots; it does not distribute the original raw RGB datasets. New demonstrations produce a new experiment, not an exact retraining of an unavailable historical dataset.

Before training

Install the pinned policy overlay, collect whole episodes and pass the task's independent dataset audit. Confirm ordered state/action fields, image processing, control rate and train/validation splits. Keep test observations out of training and checkpoint selection. See collection.

Native PushSlider example

Inspect these commands first. --dry-run prints the original backend without launching a learner:

.venv/bin/python scripts/benchmark.py train --task PushSlider --model ACT \
  --data artifacts/my-slider-data --output-dir logs/my-slider-act \
  --episodes 80 --steps 20000 --batch-size 8 --seed 42 --dry-run
.venv/bin/python scripts/benchmark.py train --task PushSlider --model DP \
  --data artifacts/my-slider-data --output-dir logs/my-slider-dp \
  --episodes 80 --epochs 40 --batch-size 64 --microbatch 8 --seed 42 --dry-run

Remove --dry-run only when intentionally starting that experiment. Output directories must be fresh. DP gradient accumulation preserves effective batch size 64; a smaller microbatch changes memory use, not that effective batch size. Initial encoder weights and their provenance must match the selected protocol.

Preserve the actual comparison

The core ACT configurations use 80 training episodes. DP uses 76 training and four held-out episodes from 80 available demonstrations. ACT and DP budgets are not equal compute. Original per-task configs in benchmarks/core/configs/ and research/core-datasets.json are authoritative, including other task-specific execution details. HotStab uses 59 successful training demonstrations and eight separate development states, with its own budgets and cameras.

Archive the data identity, source hashes, config, normalization statistics, training seed, logs and checkpoint-selection rule. Reload a checkpoint in a clean process and inspect predictions on recorded observations before a closed-loop run. Select checkpoints using the declared validation protocol, then evaluate on a separate fixed test set. Published test sets are for replication, not tuning.

EE interface study

The recorded EE experiments use prepare_research_dataset.py and the train_research_act.py / train_research_dp.py backends. Passing --interface ee to the common launcher selects those trainers for PushSlider/PullLever. An EE view must have audited expert-command replay and its own normalization. It is not an interchangeable label on a native actuator dataset. See interfaces.