Measured on the M5 Mac via Desktop Commander, 29 September 2026. This fixes control flow and evidence validation, not the learned route policy. The two-dish incumbent and experimental full-kitchen weights are unchanged.
Reproduced defects
Nine regression tests failed on the previous code before the patch:
- A no-damage trial could continue after damage or a lost life. A synthetic later health reset could make the episode report success despite its history.
- An unsafe initial state could execute controls.
- A NaN uncertainty value could pass the comparison and execute an action before evidence serialization failed. No successful historical run is asserted to have encountered this input; it was reproduced with an invalid fixture.
- Model failure before the first control lost the authoritative initial outcome.
- Boolean/fractional step limits and invalid stall/uncertainty limits were not rejected consistently before creating evidence.
- Replay proof flags were tested for truthiness instead of exact booleans; claimed goal success and verified-frame counts were not fully cross-checked.
The original red-test output is retained under docs/evidence/episode-logic/.
Fixes
runner.py freezes a separate survival contract for Pluto’s two_dishes and
kitchen1_clear objectives. It checks authoritative outcomes at action boundaries,
records the failing action and outcome, saves the terminal checkpoint, and stops
with reason=safety_violation. It issues no subsequent actor control after an
observed violation. The actor receives neither this contract nor judge memory.
The runner rejects non-finite/out-of-range uncertainty before issuing controls. It validates budgets before creating a run and preserves the initial judge result when inference fails. Other games and the legacy first-dish objective do not silently acquire the stricter Pluto survival rule.
learning.py now checks replay before-ticks as well as pixels, action macros,
after-ticks and outcomes. The final serialized emulator snapshot must match the
recorded terminal snapshot byte-for-byte, rather than merely checking that the
checkpoint file has its recorded hash.
Goal certificates must agree with the report, termination reason, final outcome, survival history and exact frame count. New runs require the terminal-state and safety certificates. Consistent historical evidence remains readable without inventing new certificates or rewriting old files. Hashes provide content integrity, not signatures proving trust in an arbitrary external author.
Real emulator regression
No training or checkpoint replacement occurred in this change. Both frozen paired protocols were rerun, with unchanged controls, starts and maximum budgets. These are reproductions, not new held-out sets.
| Check | Result |
|---|---|
| Two-dish candidate vs first-dish baseline | 10/10 vs 0/10; gate passed |
| Full-kitchen V3 vs two-dish baseline | 8/10 vs 0/10; gate still failed |
| Paired action replays and terminal snapshots | 40/40 exact |
| Direct full-kitchen launcher | 99 controls; stage 5; full health/lives |
| Direct-launch replay and neutral next-stage check | Passed |
At x=8, the failed V3 run now stops at control 92 rather than 180. At x=52 it stops at control 86 rather than 180. Both are recorded as failures, with their first observed life-loss outcome preserved. Stopping a failed trial is not a learned recovery or a new successful policy.
Validation and preserved artifacts
165 tests passed, zero skipped: 102 pixel/episode tests, 53 preserved Mario wrapper tests, and 10 game tests. This includes 12 new regression tests: the nine originally failing cases plus terminal-snapshot and safety-certificate checks. Ruff Python-error checks, formatting, launcher syntax and Git diff checks passed.
Evidence and exact source hashes are in docs/evidence/episode-logic/validation.json.
The paired summaries retain every failure and success. Direct play repeats an
already-tested scenario and must not be counted as independent evaluation.
No ROM, model weight, raw screenshot or emulator save is committed here. MOSS, Mario, the two-dish winner and the full-kitchen V3 artifact were not edited. No service, login agent or GitHub Actions workflow was added.
Remaining gameplay boundary
The full-kitchen model remains experimental at 8/10 in this protocol. The strict 10/10 promotion requirement is unchanged. Correcting the policy still requires better independently verified recovery demonstrations and a new frozen test set; this commit does not claim to have solved those remaining learned-control errors.