Back to the experimentSupporting notebook

Pluto's Corner: cumulative two-dish visual policy

Source: rgba-game-learning/docs/TWO_DISH_CURRICULUM.md · revision 6550ead3945b

Measured on the M5 Mac, 29 September 2026. Scope: the first TWO dishes in kitchen stage 1, not all four dishes, a stage clear, or a full-game win.

Outcome

The corrected candidate completed both dishes in 10/10 paired held-out trials. The frozen first-dish baseline completed only the first dish in the same trials: 0/10 two-dish successes. All candidate runs retained health 3 and lives 3, and all 20 candidate/baseline traces replayed exactly. Baseline and candidate used exactly the same emulator starting state, button catalogue, 60-action limit and 120-second wall budget. Evaluation order alternated. No incumbent checkpoint was overwritten.

Prepared x Candidate decisions Candidate dishes Baseline dishes
4 31 2 1
6 38 2 1
10 36 2 1
14 36 2 1
22 37 2 1
26 36 2 1
30 35 2 1
34 35 2 1
38 36 2 1
42 35 2 1

Positions are setup provenance, never actor input. These are nearby starts in the same deterministic task and trajectories can converge. They are not independent levels or evidence of broad game competence. Do not infer a population win rate.

Development history, including failures

  1. Source-assisted exploration found a route to two dishes preserving health and lives. Replay-verified demonstrations were collected at x=18, x=32 and x=46.
  2. Candidate v1 trained on 73 labelled frames from x=18 and x=32. Another 35 frames from x=46 selected the checkpoint. Offline label accuracy reached 100%, but actual gameplay scored only 8/10. Starts x=20 and x=40 lost lives. The fixed gate rejected it. Those results are retained, not erased or relabelled.
  3. Successful source-assisted corrections at those two failing starts added 70 training examples. Candidate v2 trained from scratch using 143 frames from four starting groups, with the same 35-frame model-selection group. It is a new cumulative policy, not a weight update to the preserved first-dish model.
  4. Before v2 evaluation, a new ten-scenario protocol was frozen. The evaluator excludes raw emulator states used by either model’s training/model selection, AND all v1 evaluation states. Goal labels cannot manufacture a new scenario.
  5. V2 passed 10/10 on that new set. A later hardened-runner repetition also passed 10/10, but it is a repeat of the same set, not ten additional held-out trials.

Selected v2 epoch: 65; training NLL 0.00009912, selection NLL 0.00064219. Both label accuracies are 100% on small source-assisted trajectories. Gameplay, not those numbers, determines the local milestone gate.

Model and actor boundary

The model reuses clm.heads.make_head: a 128-wide state projection and 64-wide action projection, each depth 2 and projection dimension 64. Inputs are two 40x36 RGB images derived from owned RGBA frames, plus three previous controls. The actor receives no RAM, OAM, positions, dish count, future frame or route index. Both dishes are played by the network; only menus and starting-position setup are scripted before scoring. No VLM, paid API or forward-search controller runs.

Training used CPU, seed 177, AdamW learning rate 0.001, weight decay 0.0001, gradient norm limit 1 and cosine logit scale 20. It stopped after 100 epochs without a selection improvement within the 600-epoch budget.

Frozen candidate SHA-256: 8929079757a58c9c78b82912377b5355272185a21957587dba7dfab9e2789d8f

Frozen first-dish baseline SHA-256: 2ca7e614a37e213ba56c24f87ef65497592f9ceb8517fd90b0f4879f434f813d

Weights and complete private traces remain on the Mac, outside Git. The local milestone gate does not modify the broader project’s promotion rules or claim that kitchen stage 2 was reached. Stage state stays 4; a stage clear requires 5.

Watch the new policy on the Mac

cd "$WORKSTATION/Projects/experimentos-pluto-20260929/rgba-game-learning"
export PIXEL_PYTHON="$WORKSTATION/Projects/mario-rgba-pokemon/pixel/.venv/bin/python"
export CLM_PYTHON="$WORKSTATION/clm-mac/src/CLM/.venv/bin/python"
bash pixel/pluto_head.sh play \
  --rom "$WORKSTATION/Games/OpenSource/plutoscorner/build/PLUTOS_CORNER.gb" \
  --spec pixel/configs/plutos_corner_two_dishes.json \
  --checkpoint "$WORKSTATION/clm-mac/models/pluto-two-dish-winner/candidate.pt" \
  --training-report "$WORKSTATION/clm-mac/models/pluto-two-dish-winner/training-report.json" \
  --shift-frames 1 --max-steps 60 --window \
  --out "$WORKSTATION/clm-mac/pluto-two-watch-$(date +%Y%m%d-%H%M%S)"

The frozen checkpoint passed a visible SDL2-window repetition at x=34: 35 scored decisions, two dishes, full health/lives, exact replay, 9.04 seconds of paced game-loop time. This repeats a scenario and is not an extra holdout.

Omit --window for accelerated headless play. The renderer closes after the trial and replay. Repeating this command is a reproduction, not a new holdout.

Reproduce training

Use the same environment exports above and a fresh output directory:

B="$WORKSTATION/clm-mac/pluto-two-retrain-$(date +%Y%m%d-%H%M%S)"
ROM="$WORKSTATION/Games/OpenSource/plutoscorner/build/PLUTOS_CORNER.gb"
SPEC=pixel/configs/plutos_corner_two_dishes.json
for shift in -8 0 -6 4 8; do
  bash pixel/pluto_head.sh demo --rom "$ROM" --spec "$SPEC" \
    --shift-frames "$shift" --out "$B/demo-$shift"
done
bash pixel/pluto_head.sh train \
  --train-runs "$B/demo--8" "$B/demo-0" "$B/demo--6" "$B/demo-4" \
  --validation-runs "$B/demo-8" --out "$B/model" --epochs 600

The demo command is explicitly source-assisted, not learned gameplay. Only successful replay-certified runs are accepted for training. This implementation trains a new model from scratch; the preserved original models are unchanged. Different dependencies/hardware need not reproduce byte-identical weights.

Paired evaluation

pixel/pluto_head.sh benchmark --protocol FILE --out NEW_DIRECTORY runs the frozen paired protocol and rejects changed model/report/ROM/verifier hashes. The checked-in v2 protocol points to the measured artifacts on this Mac:

bash pixel/pluto_head.sh benchmark \
  --protocol docs/evidence/two-dishes/v2/protocol.json \
  --out "$WORKSTATION/clm-mac/pluto-two-paired-repeat-$(date +%Y%m%d-%H%M%S)"

A reproduction needs the referenced original evidence directories, because training and prior-evaluation exclusions are verified against those artifacts. To evaluate a new model elsewhere, create a new protocol with local paths and hashes before running it. Never present reused model-selection/development scenarios as new held-out evidence.

Post-review logic hardening

A later correctness review tightened the evaluator without changing either frozen checkpoint. The local milestone gate now requires all 10/10 candidate scenarios to succeed; the earlier 90% threshold was removed because this is a small, deterministic nearby-start benchmark. The observed v2 result was already 10/10, so this makes the rule match the evidence rather than changing the outcome.

The benchmark now also requires the replay certificate’s verified_goal field to agree with run success, requires successful runs to terminate via goal_verified, checks both the raw emulator-state fingerprint and prepared x position for baseline/candidate equality, and verifies the loaded actor identity against the frozen checkpoint hash.

The ordinary play path now excludes training/model-selection raw emulator starts independently of goal labels, closing the same label-reuse loophole that the paired benchmark had already guarded against. Fresh starts also require full health before scored play begins.

The frozen v2 protocol was rerun after these changes on the same ten scenarios. This is a reproduction, not a new holdout set. It again produced candidate 10/10 versus baseline 0/10, with full health/lives and exact replay for every run. See docs/evidence/two-dishes/logic-hardening-20260929.json.

Evidence and next boundary

docs/evidence/two-dishes/ preserves v1 failures, v2 results, frozen protocols, training reports, per-run manifests/outcomes/replay certificates and the final hardening repetition. Complete screenshots and emulator states are not committed. Wall-clock game-loop timing excludes process/model startup, menu preparation and replay. Accelerated timing is not a live-display FPS or input-latency benchmark.

The next gameplay task is dish 3, then dish 4 and the state 4 -> 5 transition. This commit deliberately stops at the tested two-dish milestone. MOSS, the Mario winner and the original first-dish checkpoint are preserved. No auto-start service or paid inference endpoint was added.

Final local validation

After the review hardening, 142 tests passed with zero skips: 79 tests against this experimentos pixel/curriculum implementation, plus 53 preserved Mario wrapper tests and 10 preserved game-source tests in the original workspace. The original 138-test validation remains preserved as historical evidence. Ruff Python error checks (--select F), formatting, shell syntax and Git diff checks also passed. Validation runs on the Mac via Desktop Commander, never GitHub Actions. The two-dish benchmark was rerun after hardening and again passed 10/10 versus 0/10. Original validation details and source hashes: docs/evidence/two-dishes/validation.json. Post-review results: docs/evidence/two-dishes/logic-hardening-20260929.json.