Skip to content

Tasks and evaluation contracts

Current open-core evaluation

The frozen revision uses the following complete, ordered primary cohorts for all three training seeds of ACT, DP and BC. Every horizon counts 30 Hz control steps; OpenHatch policy targets use 20 Hz, while the other policies use 30 Hz.

Task Primary test reset IDs Control steps
PressButton 60000–60029 470
RotateValve 61000–61029 2240
OpenHatch 62000–62029 1790
CollectShell 63000–63029 7190
PushSlider 64000–64029 1790
PullLever 65000–65029 1790

The revision protocol declares separate training, validation, controller and interface streams. Follow the open-core package workflow for data and model restoration. The complete new evaluation remains in progress; historical rows cannot fill missing new-profile results.

Historical task coverage

The six-task visual core, separate HotStab, and corrected CabinRecovery/CargoRelease evaluations cover nine tasks with ACT/DP. Wall Toss has a separate paired expert demonstration, with no learned-policy score. The website catalog also includes planned environments; its size is not the number of completed model benchmarks.

Exact recorded evaluator mapping

Core task Reported test seeds Recorded control steps Original evaluator
PressButton 9110–9139 470 scripts/rollout_button_visual.py
RotateValve 2500–2529 2240 artifacts/valve_visual_20260923/act_test31/source_evaluate_valve_visual.py
OpenHatch 4400–4429 1790 scripts/evaluate_hatch_visual.py
CollectShell 5500–5529 7190 scripts/evaluate_shell_visual.py
PushSlider 8310–8339 1790 scripts/rollout_marine_visual.py
PullLever 8510–8539 1790 scripts/rollout_marine_visual.py

These are the original ACT baseline commands; the registry records both models and every study row. Some invocations also include standalone seed 42, which is excluded from the paper's 30-state denominator. Do not replace the recorded steps with a rounded nominal horizon. Valve's original evaluator is transported in core-metadata; the launcher verifies its hash before executing it.

Physical completion

Task Contract summary
PressButton Button travel with the recorded alignment, motion and contact checks
RotateValve At least 170 degrees of signed rotation; engagement/stability are diagnostics in this contract
OpenHatch Opening angle, rotation accumulated under grasp, opposing contact and stable hold
CollectShell Verified grasp/lift, transport, complete footprint inside hoop, release and settling
PushSlider Passive mechanism travel with tool engagement and base stability
PullLever Passive mechanism travel with tool engagement and base stability
HotStab Insertion, release, withdrawal and two seconds of unsupported seating
CabinRecovery Open the cabinet under grasp; extract and hold the HDD for 2 seconds
CargoRelease Sever and separate the line; open/withdraw cutter with low residual contact and speed
Wall Toss (expert only) Intentional release, free travel, complete passage and contact-free hold beyond the wall; event-based stand-off

These summaries are not replacements for versioned predicates. The exact contract travels with each original report/trajectory, config and source hash. Progress toward a threshold and complete task success are different quantities.

Which command to use

benchmark.py offers a common development entry point for collection, training and short inference checks. Exact paper replication uses paper_release.py recorded, because historical native-marine and Valve evaluators differ from some generic development routes. See reproduction.

HotStab requires its separately pinned runtime and full test IDs 3000–3029. The original split, two-camera inputs and model-selection protocol are in HotStab protocol. Its scores are excluded from the six-task average. Corrected shipwreck studies retain their own protocols and audits. Wall Toss is an expert example of current sensitivity, outside every learned-policy average.