The question
The visual policy had to act from owned frames and recent controls rather than hidden emulator state. The first-dish baseline established a small replayable milestone; the later curricula ask harder questions without rewriting that first result.
From the original notebook
Current milestone: two dishes
The cumulative RGBA-only candidate passed 10/10 new paired two-dish trials with full health/lives, compared with 0/10 for the frozen first-dish baseline. The first candidate scored 8/10 and was rejected; its failures and corrective training are retained. See two-dish findings and runbook. This is still kitchen stage 1, not a full-stage or full-game win.
Verified baseline, 29 September 2026
This experiment reuses the Mario project’s owned RGBA buffers, bounded actions, CLM projection-head architecture, immutable run evidence and deterministic replay. The small visual model is not a fine-tuned Qwen or VLM. It takes previous/current RGBA screenshots (composited and downsampled to RGB) plus its last three controls. It receives no emulator RAM, coordinates, source state or scripted runtime actions. Menus and starting positions are prepared by an explicitly unscored script.
The first-dish model was trained on 31 source-assisted demonstration rows from starts x=18 and x=32; 14 rows from x=46 were used for checkpoint selection. The selected epoch was 25 (100% training and model-selection label accuracy). These small-data accuracy numbers are not an estimate of full-game competence.
| Evaluation run | Starting x | Actions | Success | Health/lives |
|---|---|---|---|---|
| eval-n15 | 2 | 17 | First dish verified | 3/3 |
| eval-n12 | 8 | 17 | First dish verified | 3/3 |
| eval-n4 | 24 | 15 | First dish verified | 3/3 |
| eval-p4 | 40 | 15 | First dish verified | 3/3 |
| eval-p12 | 56 | 13 | First dish verified | 3/3 |
| eval-pos15 | 56 | 13 | First dish verified | 3/3 |
All six recorded runs have distinct emulator start hashes, avoid the training and model-selection start hashes, and replay exactly. They cover five horizontal start positions in one microtask, not six independent levels or broad transfer. The two x=56 starts differ in emulator state/timing. This is limited evidence of robustness to initial-position changes, not proof of general gameplay ability.
Checkpoint: 2ca7e614a37e213ba56c24f87ef65497592f9ceb8517fd90b0f4879f434f813d
(4,488,617 bytes). Weights, ROM, emulator saves, screenshots and environments stay
outside Git. Original JSON reports are in docs/evidence/first-dish.
Import hashes are in docs/import-provenance.json.
Run on the prepared M5 Mac
From this directory, use the existing isolated environments without modifying MOSS:
export PIXEL_PYTHON="$WORKSTATION/Projects/mario-rgba-pokemon/pixel/.venv/bin/python"
export CLM_PYTHON="$WORKSTATION/clm-mac/src/CLM/.venv/bin/python"
bash pixel/pluto_head.sh play \
--rom "$WORKSTATION/Games/OpenSource/plutoscorner/build/PLUTOS_CORNER.gb" \
--checkpoint "$WORKSTATION/clm-mac/pluto-head-experiment/model/candidate.pt" \
--training-report "$WORKSTATION/clm-mac/pluto-head-experiment/model/training-report.json" \
--shift-frames 4 --out "$WORKSTATION/clm-mac/pluto-check-$(date +%Y%m%d-%H%M%S)"
--window enables a paced emulator window. The full source runbook is retained
in pixel/README.md; its old host paths are historical and may
be overridden. Runtime API versions are pinned in pixel/requirements.lock.txt.
Limits and next experiment
First-dish success does not mean Kitchen Counter 1 is cleared. That next test is now documented in the two-dish runbook above. The next unsolved milestone is dish 3, then dish 4. Continue toward four dishes and the state 4 → state 5 level transition. Never promote based on offline accuracy or source-assisted replay alone. Preserve the original checkpoint and record failures as well as successes.
Evidence identifiers include timings and artifact paths; new runs need not have identical dataset hashes. A replay compares the recorded run, not wall-clock speed. The timings in JSON exclude interpreter/model startup and unscored menu setup.
Game attribution and licence
Pluto’s Corner by bbbbbr: https://github.com/bbbbbr/plutoscorner (release commit
7e362c3). The game is source-available under CC BY-NC-SA 4.0: free for this
noncommercial experiment, not unrestricted commercial open-source software.
The ROM and game assets/source are not republished here. Consult the upstream
licence before redistributing assets or using them commercially. The pipeline is
separate from the game; this experiment grants no additional rights to the game.
Post-review hardening
The two-dish evaluator was subsequently tightened to require 10/10 candidate
successes, replay-goal consistency, raw-start exclusion independent of goal
labels, and explicit baseline/candidate start equality. The frozen v2 checkpoint
still passes 10/10 versus 0/10 on the same reproduction protocol. See
docs/TWO_DISH_CURRICULUM.md.
Full-stage end-to-end experiment (not promoted)
Kitchen Counter 1 has now been cleared by an RGBA-only learned policy in 8/10 trials of one frozen protocol (two-dish baseline: 0/10). A separate richer-input candidate scored 7/10 on a different protocol. Both strict gates rejected promotion. The validated two-dish incumbent remains unchanged. Full methods, failures and watch commands: KITCHEN1_END_TO_END.
Episode control-flow fixes
The latest runtime change records and stops irreversible health/life-loss trials, rejects invalid uncertainty before controls, and verifies complete terminal checkpoints and replay certificate consistency. All 165 software tests pass; the frozen two-dish and full-kitchen results remain 10/10 and 8/10 respectively. No model was retrained or promoted. See episode logic fixes.
30 September: corrective-data full-kitchen experiment
The new corrected policy passed 10/10 full-kitchen scenarios against a matched Linux retraining control at 7/10. The complete launcher repeated candidate 10/10 with the identical candidate checkpoint; the retrained control was 5/10 on that repetition. These are one level’s nearby scenarios, not a full-game win. The Mac went offline before training, so this is CPU-only Ubuntu evidence and Mac validation/deployment is pending. Original Mac checkpoints remain unchanged. See corrective learning for the frozen comparison, source corrections, limitations and single-command reproduction.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.