Skip to content

Execution sensitivity of the reported conclusions

Same final checkpoints, three training identities and 30 task resets on two workstations. These paired repetitions do not replace primary scores.

Model Primary mean (%) Repeat mean (%) Drift (pp)
ACT 43.5 43.1 -0.4
DP 61.9 60.4 -1.5
BC 26.3 24.4 -1.9

The equal-task ordering DP > ACT > BC is preserved. Across the 18 task-level pairwise model contrasts, no nonzero difference reverses direction. The ACT/DP tie on RotateValve becomes a one-episode DP advantage across 90 episodes; this is not evidence of a resolved difference.

Task Model Primary /90 Repeat /90 Drift (pp) Matching /90
PressButton ACT 41 46 +5.6 65
PressButton DP 54 50 -4.4 56
PressButton BC 36 32 -4.4 68
RotateValve ACT 76 75 -1.1 85
RotateValve DP 76 76 +0.0 90
RotateValve BC 67 63 -4.4 86
OpenHatch ACT 73 76 +3.3 77
OpenHatch DP 90 90 +0.0 90
OpenHatch BC 15 14 -1.1 71
CollectShell ACT 0 0 +0.0 90
CollectShell DP 30 24 -6.7 62
CollectShell BC 0 0 +0.0 90
PushSlider ACT 9 7 -2.2 84
PushSlider DP 38 40 +2.2 88
PushSlider BC 2 2 +0.0 88
PullLever ACT 36 29 -7.8 77
PullLever DP 46 46 +0.0 88
PullLever BC 22 21 -1.1 83

Outcome agreement is 1,438/1,620 (88.8%). Small aggregate drift can conceal opposing episode flips. The largest absolute task/model mean drift is 7.8 pp (PullLever ACT).

The separate same-GPU seed-17 DP repeat agrees on 169/180 outcomes (93.9%); task counts change by at most one success out of 30, despite seven Shell and three Slider outcome flips. This does not identify the numerical source of execution variability. Controller, action-interface and current contrasts were not rerun across workstations in this comparison.

Regenerate: python scripts/summarize_execution_variation.py. Source hashes and all 18 comparisons: research/revision-v2-execution-impact.json.