Back to the experimentSupporting notebook

Episode termination and replay integrity fixes

Source: rgba-game-learning/docs/EPISODE_LOGIC_FIXES.md · revision 6550ead3945b

Measured on the M5 Mac via Desktop Commander, 29 September 2026. This fixes control flow and evidence validation, not the learned route policy. The two-dish incumbent and experimental full-kitchen weights are unchanged.

Reproduced defects

Nine regression tests failed on the previous code before the patch:

The original red-test output is retained under docs/evidence/episode-logic/.

Fixes

runner.py freezes a separate survival contract for Pluto’s two_dishes and kitchen1_clear objectives. It checks authoritative outcomes at action boundaries, records the failing action and outcome, saves the terminal checkpoint, and stops with reason=safety_violation. It issues no subsequent actor control after an observed violation. The actor receives neither this contract nor judge memory.

The runner rejects non-finite/out-of-range uncertainty before issuing controls. It validates budgets before creating a run and preserves the initial judge result when inference fails. Other games and the legacy first-dish objective do not silently acquire the stricter Pluto survival rule.

learning.py now checks replay before-ticks as well as pixels, action macros, after-ticks and outcomes. The final serialized emulator snapshot must match the recorded terminal snapshot byte-for-byte, rather than merely checking that the checkpoint file has its recorded hash.

Goal certificates must agree with the report, termination reason, final outcome, survival history and exact frame count. New runs require the terminal-state and safety certificates. Consistent historical evidence remains readable without inventing new certificates or rewriting old files. Hashes provide content integrity, not signatures proving trust in an arbitrary external author.

Real emulator regression

No training or checkpoint replacement occurred in this change. Both frozen paired protocols were rerun, with unchanged controls, starts and maximum budgets. These are reproductions, not new held-out sets.

Check Result
Two-dish candidate vs first-dish baseline 10/10 vs 0/10; gate passed
Full-kitchen V3 vs two-dish baseline 8/10 vs 0/10; gate still failed
Paired action replays and terminal snapshots 40/40 exact
Direct full-kitchen launcher 99 controls; stage 5; full health/lives
Direct-launch replay and neutral next-stage check Passed

At x=8, the failed V3 run now stops at control 92 rather than 180. At x=52 it stops at control 86 rather than 180. Both are recorded as failures, with their first observed life-loss outcome preserved. Stopping a failed trial is not a learned recovery or a new successful policy.

Validation and preserved artifacts

165 tests passed, zero skipped: 102 pixel/episode tests, 53 preserved Mario wrapper tests, and 10 game tests. This includes 12 new regression tests: the nine originally failing cases plus terminal-snapshot and safety-certificate checks. Ruff Python-error checks, formatting, launcher syntax and Git diff checks passed.

Evidence and exact source hashes are in docs/evidence/episode-logic/validation.json. The paired summaries retain every failure and success. Direct play repeats an already-tested scenario and must not be counted as independent evaluation.

No ROM, model weight, raw screenshot or emulator save is committed here. MOSS, Mario, the two-dish winner and the full-kitchen V3 artifact were not edited. No service, login agent or GitHub Actions workflow was added.

Remaining gameplay boundary

The full-kitchen model remains experimental at 8/10 in this protocol. The strict 10/10 promotion requirement is unchanged. Correcting the policy still requires better independently verified recovery demonstrations and a new frozen test set; this commit does not claim to have solved those remaining learned-control errors.