Collect and audit demonstrations¶
Use a development checkout and new reset states. Do not modify a pinned runtime or tune against the published test cohort. A successful expert establishes feasibility under its contract, not learned-policy success.
The examples below describe the historical collection tools. For current
open-core primary data, follow the fixed candidate streams and
audit_revision_dataset.py in the revision protocol,
or run its complete source campaign. The historical
dataset_audit.json and revision revision_audit.json are different contracts.
Record a first development episode¶
After installation, run the PressButton collection example in
Verify installation. Full primary datasets use the frozen
candidate streams in run_revision_campaign.py; arbitrary development episodes
are not primary training data.
Record one native marine episode¶
PushSlider and PullLever share the native 30 Hz marine recording schema. This example records one development episode and its wrist RGB:
OMP_NUM_THREADS=4 .venv/bin/python scripts/rollout_marine_visual.py \
--task PushSlider --mode collect --purpose development --seeds 12000 \
--steps 1790 --output-dir artifacts/my-slider-pilot
Collection uses privileged expert feedback. Deployment of a visual policy must
use --mode policy; it must not call the expert for replacement actions.
Retain the report, failed episodes, source hashes and physical trace. Inspect
approach, grasp, mechanism motion and release in their recorded order.
Audit the pilot¶
.venv/bin/python scripts/audit_marine_dataset.py \
--data artifacts/my-slider-pilot --episodes 1 --pilot
The audit requires a successful whole episode. It checks RGB, frame/action alignment, finite state/commands, actuator packing and physical success replay. A failed collection episode is not made successful by changing the predicate. A pilot audit is deliberately not accepted by the final trainers.
Build a training dataset¶
Collect batches into a dedicated dataset directory using unique, declared training seeds, separate from validation and test resets:
my-slider-data/
train_batch00/seed_<id>/metadata.json
train_batch00/seed_<id>/trajectory.npz
train_batch00/seed_<id>/wrist.rgb
train_batch01/...
Once 80 successful demonstrations are available, audit the complete selected set:
.venv/bin/python scripts/audit_marine_dataset.py \
--data artifacts/my-slider-data --episodes 80
The resulting dataset_audit.json identifies the exact ordered successful seeds
and file hashes. The trainers reject a mismatched task, insufficient data, a pilot
report or a different audited selection. Existing audit files are preserved.
Raw RGB can be large: inventory the requested episodes before transferring data.
Other protocols¶
RotateValve, OpenHatch, CollectShell and PressButton retain their original recorders and loaders. Use the backend mapping in task reference and the task's frozen protocol. HotStab uses its isolated runtime, two robot cameras and a geometric audit. Do not convert datasets by renaming files or assume that all six tasks share the same action vector. See data contracts.