Measured on the Apple M5 Mac on 29 September 2026 through Desktop Commander. The model has demonstrated full first-kitchen clears, but neither tested full-stage candidate passed the strict 10/10 promotion gate. The validated two-dish incumbent is unchanged. This is not a full-game victory.
Measured outcomes
| Candidate | Inputs / data | Scored result | Gate |
|---|---|---|---|
| v1 | Two 40x36 frames, 3 controls; 296 training / 98 selection rows | 0/3 development clears | Not promoted |
| v2 | Two 64x58 frames, 8 controls; same data | 0/3 development clears | Not promoted |
| v3 | Two 64x58 frames, 8 controls; 1,277 training / 199 selection rows | 8/10 full-stage clears; baseline 0/10 | Rejected: two runs lost a life |
| v4 | Two 80x72 frames, 12 controls; same expanded data | 7/10 full-stage clears; baseline 0/10 | Rejected: one life-loss run and two stalls |
V3 and v4 used DIFFERENT frozen evaluation sets. Their percentages are not a controlled architecture ranking. Both used the frozen two-dish baseline on identical starting states and budgets within their own paired protocol. V4 excluded all v3 evaluation states. No lower promotion threshold was substituted.
The v3 failures began at x=8 and x=52. In v4, x=10 reached stage 5 after losing a life and therefore correctly did NOT count as a success; x=30 and x=34 stalled at three dishes. Successful candidate trials preserved full health and all three lives. All completed paired traces replayed exactly.
What end-to-end means here
Only title/intro and starting-position/timing setup are scripted before scoring. The scored actor controls the complete route, including dishes 1 and 2, from the start of kitchen 1. It sees current/previous RGBA-derived pixels and recent executed controls, not RAM, sprite coordinates, dish counts, a route index, future frames or a VLM’s answer. It reuses actual CLM projection heads; this is not a fine-tuned Qwen model. Teacher generation is separately source-assisted.
Each baseline/candidate pair uses the same emulator state, action catalogue, 180-action and 120-second budget. Raw start hashes are checked against both models’ training/selection sources and the frozen development ledger, regardless of goal labels. Scenarios are nearby position/timing variations of one level; they are not independent levels and do not support a population win-rate claim.
A clear requires stage 4 -> 5, health 3 and lives 3. The gate additionally requires no damage/life loss anywhere in the trace and exact replay. After replay, a separate unscored 120-frame neutral check confirms kitchen 2 is still active and healthy. This settling check is NOT extra model-controlled gameplay.
Verifier boundary bug found and fixed
The dish counter at 0xC3A8 belongs to a kitchen-1 sprite. Kitchen 2 reuses that memory: a real transition produced the byte 13, not a valid dish count. The old cross-stage plausibility check would incorrectly reject this transition. The verifier now validates the dish count only in kitchen 1 and reports null outside it. It never presents reused sprite memory as a count of completed dishes. Four dishes without the stage transition are not sufficient for stage-clear credit.
Training, failures and provenance
A source-assisted route first established a safe complete stage traversal. Fixed button counts were fragile at the lower counter and third-dish edge. The recorded source sweeps retain failures rather than relabelling them as successful teaching. A feedback teacher uses source-defined sprite information only when constructing demonstrations, restores the exact initial state, and records/replays the resulting bounded controls. The evaluated learned actors never call that teacher.
V3/v4 trained from 13 successful starting groups and used 2 other groups for checkpoint selection. Both achieved 100% offline action recall, which did not prevent gameplay failures. The expanded training set had 1,277 labelled frames; selection had 199. Failed correction-source attempts were not added to training. The shipped teacher has a bounded final-gap search; it is not an evaluated actor.
Earlier attempted stage benchmarks failed a 121-frame settling-action validation before completing the matrix. Those partial attempts are retained. The corrected check uses 119 held-neutral frames plus 1 release frame, respecting the 120-frame limit. Completed benchmark results, not partial attempts, supply the denominators.
V3’s frozen experimental checkpoint is 11,483,561 bytes, SHA-256:
3a9dfab6da5808fdae9ac4a15be92300e9f3126ab69b96f35ea3758e70184e11.
V4 is 17,787,305 bytes, SHA-256:
3e8f1b2a896debc07e95b2b833c8e0a5447a6a7713d8f6153177a5279c356b2b.
The old pixel recipes and saved models retain their original loading contracts.
Raw screenshots, cartridge copies, weights and emulator states stay outside Git.
See docs/evidence/kitchen1/ for model/benchmark/source summaries and certificates.
Watch on the prepared Mac
cd "$WORKSTATION/Projects/experimentos-pluto-20260929/rgba-game-learning"
bash pixel/watch_kitchen1.command
This uses the frozen v3 experimental artifact, setup shift 0, four waiting frames, then the actual policy with an SDL2 window. It is a reproduction of a known successful scenario, not a new holdout evaluation. It also runs exact replay and checks the next-stage transition. MOSS, Mario and the two-dish incumbent remain untouched. No inference service or login agent is installed.
For headless execution use pixel/pluto_head.sh play with
--spec pixel/configs/plutos_corner_kitchen1.json, the experimental model/report,
--setup-wait-frames 4 --shift-frames 0 --max-steps 180 and no --window.
Versioned features are selected at training with --recipe; old checkpoints
retain the old default recipe. Checkpoint hashes and game/action-spec matching
remain mandatory.
Remaining boundary
Neither stage candidate is promoted. The validated two-dish checkpoint stays the incumbent. The next research target is safer recovery at the third/fourth-dish transition and richer independently verified correction trajectories, followed by a genuinely new frozen evaluation set. Do not repeatedly tune on this test set and then describe it as held out. End-to-end functional execution is demonstrated; reliable all-scenario stage completion remains open in this experiment.
Final verification and visible run
153 software tests passed with zero skips: 90 pixel/stage tests (including real training/loading of both new recipes), 53 preserved Mario tests and 10 game tests. Ruff Python error checks, formatting, both launcher syntax checks and Git diff checks passed. The old two-dish frozen benchmark still passed 10/10 versus 0/10. A full-stage v3 regression reproduced 8/10 versus 0/10 and correctly failed its strict gate. Regression repetitions do not increase the independent trial count.
The visible SDL2 test completed 99 scored decisions, stage 5, health 3, lives 3, exact replay and a healthy settled next stage. Its paced game-loop time was 33.91 seconds, excluding setup and replay; this is not a speed benchmark. A local MP4 reconstructs that run from decision-boundary screenshots and was opened in QuickTime. It is a recording, not a live session or a continuous-frame capture.
Evidence: visible-evaluation-context.json, visible-replay.json,
visible-stage-transition-check.json and validation.json under
docs/evidence/kitchen1/. All tests ran locally on the Mac; no GitHub Actions.
Subsequent episode-logic corrections
See episode logic fixes for the later fail-fast survival contract and stricter replay/terminal-state validation. Those changes preserve this historical experiment and do not promote the 8/10 full-kitchen candidate.