Journal

Corrective data meets a fresh gameplay gate

On Ubuntu CPU, the corrected policy passed 10/10 full-kitchen trials against a same-host retraining control at 7/10.

Vlad / experimentos.
The boundary

The repeated control reached 5/10 with different weights. This is one training seed and one level; no Mac deployment or deterministic control retraining is claimed.

Measurements

Recorded result

Corrective data vs a same-host retraining control

Ubuntu CPU · initial paired full-kitchen evaluation

Corrective data vs a same-host retraining controlRetraining control: 7 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.Retraining control7Retraining control: 7 successful trials / 10Candidate10Candidate: 10 successful trials / 10Gate: 100successful trials / 10
  1. Retraining control7
  2. Candidate10

successful trials / 10 · gate: 10

One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.

View data & source
Corrective data vs a same-host retraining control · successful trials / 10
ConfigurationValue
Retraining control7
Candidate10

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The repeat preserved the candidate result

Ubuntu CPU · repeated set · newly retrained control

The repeat preserved the candidate resultRetraining control: 5 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. The control had different weights and reached 5/10. Repeating a set adds no new holdout trials.Retraining control5Retraining control: 5 successful trials / 10Candidate10Candidate: 10 successful trials / 10Gate: 100successful trials / 10
  1. Retraining control5
  2. Candidate10

successful trials / 10 · gate: 10

One level and nearby starts, not broad game competence. The control had different weights and reached 5/10. Repeating a set adds no new holdout trials.

View data & source
The repeat preserved the candidate result · successful trials / 10
ConfigurationValue
Retraining control5
Candidate10

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

Corrective trajectories were added to a fresh same-host training comparison on Ubuntu CPU. The initial and repeated results stay separate, including the control’s changed weights, so the result is not presented as a direct improvement over the earlier Mac experiment.

From the original notebook

Measured 30 September 2026. This experiment ran on the user’s Ubuntu 12700 host, CPU only, after the Mac became unreachable through both Desktop Commander and SSH. It is NOT a new M5 measurement or a deployment to the Mac.

Result and comparison boundary

The corrected student passed 10/10 full Kitchen Counter 1 trials, versus 7/10 for a matched same-host retraining control. All ten candidate trajectories preserved health 3 and lives 3, replayed exactly (including terminal snapshots), and passed the separate neutral next-stage readiness check. The strict gate passed. No existing checkpoint was overwritten or automatically replaced.

The control was freshly trained on Linux from 13 successful source-start groups. It is NOT the original Mac V3 checkpoint. Consequently, this result must not be presented as a direct improvement from the Mac’s historical 8/10 to 10/10. Both new models use identical code, architecture, seed, optimizer, action catalogue, model-selection starts, and maximum training/evaluation budgets. The controlled difference is eight additional verified correction trajectories.

Property Control Corrected candidate
Training start groups 13 21
Training labelled frames 1,277 2,057
Selection groups / frames 2 / 199 2 / 199
Full-stage successes 7/10 10/10
Independent model-weight seed repetitions 1 1

These are nearby position/timing variations of ONE level, not independent games. They do not establish a population win rate, robustness to arbitrary disturbances, or a full-game victory. A single training seed is not an across-seed estimate.

What changed

pluto_correction.py adds a bounded, ROM-hash-bound source teacher. The old teacher had no search around the third-dish crossing and could lose a life before it could supply a correction. The new teacher searches both that crossing and the last gap using a fixed catalogue, a 20-second wall budget and 320-branch cap. Every proposed route is bounded to 180 actions and rejects health or life loss. The initial emulator snapshot is restored even after a failed branch, timeout, or emulator exception. Independent recording and exact replay must then pass before any proposed demonstration becomes training data.

The new teacher solved the two previously failing V3 starting states at x=8 and x=52 (four-frame setup wait), in 98 and 96 actions respectively. Those are source-assisted demonstrations, NOT independent student successes.

Eight correction demonstrations add 780 labelled frames. Their setup pairs (shift, wait) are (-12,4), (10,4), (-11,7), (-1,7), (1,7), (-12,2), (10,2), (-1,2). These were chosen from prior failure/development contexts and are training data.

The learner itself remains the existing CLM projection-head architecture: two 64x58 RGB-composited frames plus eight executed controls. Seed 177, AdamW learning rate 0.001, and an 800-epoch upper limit are unchanged. No Qwen model, VLM server, RAM input, planner, or teacher is invoked by the scored student. The survival contract and 10/10 gate were not relaxed.

pluto_experiment.py provides a complete collection, training and paired-test entrypoint. demo --source-search exposes the teacher explicitly; play does not accept that flag. The old default source teacher and saved models remain.

Frozen evaluation

The scenario choices were written before either new model was trained. Model and report hashes were then frozen before running the paired benchmark. Evaluation uses shifts [-14,-12,-10,-6,-2,0,2,4,10,14] with an eleven-frame setup wait, 180-action and 120-second budgets. Raw emulator-start hashes are checked against both models’ training/selection sources and the committed historical development/V3/V4 ledgers, independently of goal labels. Baseline and candidate must share the exact starting state. Both full routes, including dishes 1/2, are scored; only menus and initial-position/timing preparation are unscored.

The control failed at x=8,28,36; the candidate completed all ten scenarios. Three favorable paired differences in one small local test do not constitute strong statistical evidence across games or training seeds. No claim of statistical significance is made. The old Mac checkpoints remain the incumbents.

Both models achieved 100% action accuracy on their training/selection data. Only the separate emulator test establishes their different gameplay outcomes. The candidate checkpoint SHA-256 is: f52b89d2e992b75cc3a3b73e328f738228f2f5832a0b203581eb318cbe6e4d37. The new control SHA-256 is: d0701ab44ab240d816cfe756fe9ca1bc5ae667e6d366138d3f517299101a1820.

Reproduce on a prepared host

From rgba-game-learning, with a compatible CLM/PyTorch and pixel/PyBoy runtime:

export CLM_PYTHON=/path/to/clm-env/bin/python
export PIXEL_PYTHON=/path/to/pixel-env/bin/python
export CUDA_VISIBLE_DEVICES=''
bash pixel/pluto_head.sh corrections \
  --rom /path/to/PLUTOS_CORNER.gb \
  --out /path/to/new-correction-experiment \
  --epochs 800

The command refuses existing output paths. Failed source attempts are retained and excluded, not silently treated as wins. At least eight control-start groups, two selection groups and two correction groups must succeed to train. The actual measured run accepted all 23 planned demonstrations. A failed gate returns exit code 3; it does not overwrite an incumbent or change the pass threshold.

The measured runtime used Python 3.12, PyTorch 2.8.0+cpu and PyBoy 2.7.0. CPU training uses four Torch threads; CUDA was disabled and no other jobs were stopped. Dependency locks and upstream source commits are saved with the evidence. Different runtimes may produce different weight hashes. Running these same scenario choices again is a reproduction, not a fresh held-out evaluation.

The frozen candidate and raw evidence remain on Ubuntu under: $WORKSTATION/experiments/pluto-runtime-20260930/controlled-v1/. The repository contains code and reports, not cartridge images, model weights, raw screenshots, or emulator snapshots. Re-running training regenerates those artifacts; merely cloning the repository does not install a trained checkpoint.

Mac status

The Mac was available for the initial read-only status/hash check, then became unreachable before the experiment could start. New model training and evaluation were therefore performed on Ubuntu. No new checkpoint was installed on the Mac, and no new Mac-visible run is claimed. Mac loading and GUI validation remain pending reconnection. MOSS and the original Mac policy checkpoints were not edited.

End-to-end launcher validation and limitations

112 pixel/learning tests passed with zero skips on Ubuntu. The ten added tests cover search restoration on failure/timeouts, branch limits, invalid budgets, wrong ROMs, unsafe starts, and experiment orchestration. Python-error lint, formatting and shell syntax checks passed. The older Mac Mario-specific suites were not rerun while that machine was offline.

The first new launcher attempt stopped after collecting demonstrations because its local demo-argument variable shadowed the experiment’s options. This was fixed and regression-tested; the failed attempt is preserved, not counted as a completed training/evaluation run. The corrected single-command launcher then completed collection, both trainings, and the full paired benchmark.

That complete repetition reproduced the candidate checkpoint byte-for-byte and again achieved 10/10. Its newly retrained control achieved 5/10 rather than 7/10. All teacher action/image triples and evaluation raw-start fingerprints matched between completed runs, but the control’s weights were not byte-identical despite the same nominal seed. The numerical cause has not been isolated; deterministic retraining is NOT claimed. Both control results are preserved. The repetition is not additional independent holdout evidence or a new architecture comparison.

All 40 baseline/candidate traces across the two completed matrices had exact replay and terminal-state certificates. Candidate success cases also passed the neutral next-stage readiness check. See docs/evidence/corrective-learning/ for frozen plans, model reports, per-run certificates, dependency locks, source hashes, the failed launcher attempt, and validation metadata.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS