Journal

Learning to play from pixels

The first-dish baseline produced six successful recorded evaluations across five starting positions, with exact replay.

Vlad / experimentos.
The boundary

The later two-dish milestone is a separate frozen evaluation. Neither milestone establishes broad game competence.

The question

The visual policy had to act from owned frames and recent controls rather than hidden emulator state. The first-dish baseline established a small replayable milestone; the later curricula ask harder questions without rewriting that first result.

From the original notebook

Current milestone: two dishes

The cumulative RGBA-only candidate passed 10/10 new paired two-dish trials with full health/lives, compared with 0/10 for the frozen first-dish baseline. The first candidate scored 8/10 and was rejected; its failures and corrective training are retained. See two-dish findings and runbook. This is still kitchen stage 1, not a full-stage or full-game win.

Verified baseline, 29 September 2026

This experiment reuses the Mario project’s owned RGBA buffers, bounded actions, CLM projection-head architecture, immutable run evidence and deterministic replay. The small visual model is not a fine-tuned Qwen or VLM. It takes previous/current RGBA screenshots (composited and downsampled to RGB) plus its last three controls. It receives no emulator RAM, coordinates, source state or scripted runtime actions. Menus and starting positions are prepared by an explicitly unscored script.

The first-dish model was trained on 31 source-assisted demonstration rows from starts x=18 and x=32; 14 rows from x=46 were used for checkpoint selection. The selected epoch was 25 (100% training and model-selection label accuracy). These small-data accuracy numbers are not an estimate of full-game competence.

Evaluation run Starting x Actions Success Health/lives
eval-n15 2 17 First dish verified 3/3
eval-n12 8 17 First dish verified 3/3
eval-n4 24 15 First dish verified 3/3
eval-p4 40 15 First dish verified 3/3
eval-p12 56 13 First dish verified 3/3
eval-pos15 56 13 First dish verified 3/3

All six recorded runs have distinct emulator start hashes, avoid the training and model-selection start hashes, and replay exactly. They cover five horizontal start positions in one microtask, not six independent levels or broad transfer. The two x=56 starts differ in emulator state/timing. This is limited evidence of robustness to initial-position changes, not proof of general gameplay ability.

Checkpoint: 2ca7e614a37e213ba56c24f87ef65497592f9ceb8517fd90b0f4879f434f813d (4,488,617 bytes). Weights, ROM, emulator saves, screenshots and environments stay outside Git. Original JSON reports are in docs/evidence/first-dish. Import hashes are in docs/import-provenance.json.

Run on the prepared M5 Mac

From this directory, use the existing isolated environments without modifying MOSS:

export PIXEL_PYTHON="$WORKSTATION/Projects/mario-rgba-pokemon/pixel/.venv/bin/python"
export CLM_PYTHON="$WORKSTATION/clm-mac/src/CLM/.venv/bin/python"
bash pixel/pluto_head.sh play \
  --rom "$WORKSTATION/Games/OpenSource/plutoscorner/build/PLUTOS_CORNER.gb" \
  --checkpoint "$WORKSTATION/clm-mac/pluto-head-experiment/model/candidate.pt" \
  --training-report "$WORKSTATION/clm-mac/pluto-head-experiment/model/training-report.json" \
  --shift-frames 4 --out "$WORKSTATION/clm-mac/pluto-check-$(date +%Y%m%d-%H%M%S)"

--window enables a paced emulator window. The full source runbook is retained in pixel/README.md; its old host paths are historical and may be overridden. Runtime API versions are pinned in pixel/requirements.lock.txt.

Limits and next experiment

First-dish success does not mean Kitchen Counter 1 is cleared. That next test is now documented in the two-dish runbook above. The next unsolved milestone is dish 3, then dish 4. Continue toward four dishes and the state 4 → state 5 level transition. Never promote based on offline accuracy or source-assisted replay alone. Preserve the original checkpoint and record failures as well as successes.

Evidence identifiers include timings and artifact paths; new runs need not have identical dataset hashes. A replay compares the recorded run, not wall-clock speed. The timings in JSON exclude interpreter/model startup and unscored menu setup.

Game attribution and licence

Pluto’s Corner by bbbbbr: https://github.com/bbbbbr/plutoscorner (release commit 7e362c3). The game is source-available under CC BY-NC-SA 4.0: free for this noncommercial experiment, not unrestricted commercial open-source software. The ROM and game assets/source are not republished here. Consult the upstream licence before redistributing assets or using them commercially. The pipeline is separate from the game; this experiment grants no additional rights to the game.

Post-review hardening

The two-dish evaluator was subsequently tightened to require 10/10 candidate successes, replay-goal consistency, raw-start exclusion independent of goal labels, and explicit baseline/candidate start equality. The frozen v2 checkpoint still passes 10/10 versus 0/10 on the same reproduction protocol. See docs/TWO_DISH_CURRICULUM.md.

Full-stage end-to-end experiment (not promoted)

Kitchen Counter 1 has now been cleared by an RGBA-only learned policy in 8/10 trials of one frozen protocol (two-dish baseline: 0/10). A separate richer-input candidate scored 7/10 on a different protocol. Both strict gates rejected promotion. The validated two-dish incumbent remains unchanged. Full methods, failures and watch commands: KITCHEN1_END_TO_END.

Episode control-flow fixes

The latest runtime change records and stops irreversible health/life-loss trials, rejects invalid uncertainty before controls, and verifies complete terminal checkpoints and replay certificate consistency. All 165 software tests pass; the frozen two-dish and full-kitchen results remain 10/10 and 8/10 respectively. No model was retrained or promoted. See episode logic fixes.

30 September: corrective-data full-kitchen experiment

The new corrected policy passed 10/10 full-kitchen scenarios against a matched Linux retraining control at 7/10. The complete launcher repeated candidate 10/10 with the identical candidate checkpoint; the retrained control was 5/10 on that repetition. These are one level’s nearby scenarios, not a full-game win. The Mac went offline before training, so this is CPU-only Ubuntu evidence and Mac validation/deployment is pending. Original Mac checkpoints remain unchanged. See corrective learning for the frozen comparison, source corrections, limitations and single-command reproduction.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS