Tasks, studies and limits¶
This page records historical CAD-profile studies. For the current six-task open core, use the frozen revision protocol and reproduction status. Historical numbers below are not results of the new asset version or additional independent training runs.
The README reports nine tasks with ACT/DP evaluations: the original six-task visual core, separate HotStab, and corrected CabinRecovery/CargoRelease protocols. Only the six core rows enter the core macro-average. Core snapshots are in benchmarks/core; HotStab has its own protocol, results and per-state records.
For each core row, research/paper-results-index.json provides the exact test
states, checkpoint, horizon, evaluator hash and source report. Model configurations
in benchmarks/core/configs/ preserve input dimensions, rates, action chunks,
training budget, preprocessing and splits. The dataset inventory records that
available demonstrations and gradient-training demonstrations are not always the
same count. Control, interface and demonstration-budget comparisons have separate
cohorts and must not overwrite the suite's original scores.
| Study | Frozen snapshot |
|---|---|
| Core ACT/DP suite | benchmarks/core/model-benchmarks.json |
| EE targets vs actuator targets; DP data budgets | benchmarks/core/research-core-studies.json |
| Currents and motor capacity | benchmarks/core/water-motor-study.json |
| PID vs PD, OpenHatch DP | benchmarks/core/policy-control-study.json |
| Articulated robot systems, 30/30 each | benchmarks/articulated-embodiments-v3/summary.json |
| Custom twin-arm RotateValve, three scripted modes | benchmarks/bimanual-valve-v1/summary.json |
| Historical held-arm robot systems | benchmarks/core/embodiment-study.json |
| Separate OpenHatch ACT recovery correction | benchmarks/core/hatch-act-correction.json |
| HotStab expert / ACT / DP | benchmarks/hotstab-v1/results/ |
| Corrected CabinRecovery / CargoRelease ACT/DP | benchmarks/wreck-corrected-v3/ |
| Wall Toss, one expert-command pair | benchmarks/wall-toss-v2/ |
The held-arm comparison concerns native robot/controller systems, not isolated morphology or learned-policy transfer. Water-appearance demonstrations illustrate optical variations; they are not an ACT/DP water-domain benchmark. State PPO has its own task/observation protocol and is not an extra row in the RGB comparison.
Low success is retained as a result. HotStab ACT frequently fails to release after insertion; DP predominantly fails to establish a verified grasp. Those outcome categories are diagnostic observations, not proven unique causal explanations. The geometric audits use sampled control frames and declared conservative proxies; no claim is made about all substeps or every self-collision.
New tasks under active development are outside this frozen evaluation release. The registered source may contain additional environments, but registration alone does not make them a completed benchmark or a published result.
Repeatability and audit scope¶
OpenHatch ACT scored 10/30 in the original core evaluation and 16/30 in an instrumented repeat. Known bf16/cuDNN sensitivity does not fully explain the rollout differences. Both observations remain part of the record; the separate recovery comparison uses its own paired test cohort.
HotStab's expert scored 28/30. All 90 expert/ACT/DP test episodes passed the protocol's sampled geometric audit. Passing that audit does not establish the absence of all substep contacts or self-collisions. Core and HotStab results use one training seed per model/configuration; reset uncertainty is not training-seed variance. Hydrodynamic assumptions and illustrative water optics have not been calibrated against a physical robot.
Validated articulated PressButton comparison¶
BlueROV2/Reach Alpha and RexROV2/Oberon7 each score 30/30 on new paired resets, with the original physical success contract. Independent trace replay agrees; both motors-disabled controls fail. Configuration was frozen before opening the final states. Both systems pass 240/480 Hz contact agreement checks.
The v3 report discloses the controller fix, explicit hydrodynamic variants, reproduction commands and limits. Historical v2 failures and v1 scores are retained.
Custom two-arm RotateValve comparison¶
All three modes score 30/30 on paired final resets: left arm parked, left hand on a fixed support, and both hands turning the wheel. During turning, cooperative control has lower mean base-orientation RMS (0.060° versus 0.093°/0.095°) and mean peak motor utilization. The support grasp does not improve those averages in this calm-water setup. These scripted experts use a separate 506 mm fixture; their scores are not pooled with the original visual-policy task.
The results and audit scope retain the original world-coordinate precision failures and explain the supplemental verification. Reproduction commands use an independent runtime pin, preserving all earlier model results.