Back to the experimentSupporting notebook

RGBA game-learning pipeline

Source: rgba-game-learning/pixel/README.md · revision 6550ead3945b

Primary runnable target: Pluto’s Corner. The free/homebrew Game Boy platformer now has a ROM-hash-bound verifier and a real exact-replayed first-dish smoke test. See ../docs/PLUTOS_CORNER.md or run:

pixel/.venv/bin/rgba-game pluto-smoke \
  --rom "$WORKSTATION/Games/OpenSource/plutoscorner/build/PLUTOS_CORNER.gb" \
  --out "$WORKSTATION/clm-mac/pluto-smoke-$(date +%Y%m%d-%H%M%S)"

The Pokémon Blue adapter below remains available for a future user-owned ROM, but it is no longer required to exercise the reusable pipeline on a real game.

This is an additive extension of the Mario project, not a replacement for its validated World 1-1 checkpoint or its stricter full-route promotion gates.

Implemented and tested on the M5 Mac: owned RGBA frames, bounded controller macros, PyBoy save/restore, exact action replay, hash-linked evidence, persistent memory/checkpoints, image-chat and CLM protocol clients, grouped teacher-data export, CLM projection-head training and an independent-scenario evaluation gate.

Not demonstrated: Pokémon Blue gameplay, Pokémon learning, a complete Pokémon win, or the accuracy of a real vision-language model. No Blue ROM was found in the checked Mac project/download/game directories. No Nintendo ROM is provided. The emulator test uses PyBoy’s bundled homebrew test cartridge. The training integration test uses synthetic 8-dimensional embeddings, not Qwen or Pokémon.

What transfers from Mario

Reuse the frozen text embedding service, clm.schema.build_pairs, clm.embedder.Embedder, clm.heads.make_head, choice-head training, and the pattern of recording evidence and verifying execution. Do NOT reuse Mario’s seven-action winner weights as if they were a Pokémon policy. Mario’s recorded observations were structured state extracted from emulator RAM; that was not pixel-only play.

The new player receives only RGBA images, visibly grounded scene descriptions, recent actions and its subgoal. An optional read-only, ROM-hash-bound outcome judge can inspect milestone flags. Its data never enters the actor state. Without that judge, play can be recorded, but a victory cannot be automatically certified or exported as a successful demonstration.

Data flow

Game adapter: capture H×W×4 uint8 RGBA + execute bounded controls
    -> immutable frames, explicit alpha compositing
    -> image-capable VLM: scene/dialogue/menu/battle description
    -> bounded history + persistent subgoal
    -> visual teacher OR CLM action selector
    -> game-specific press/hold/release macro
    -> actual next RGBA frame + separate read-only outcome judge
    -> hash-linked evidence + checkpoint + optional failure gate
    -> deterministic replay proof
    -> verified teacher examples, split by complete starting-state groups
    -> new CLM state/action heads; base checkpoint preserved
    -> independent paired scenario gate; no automatic overwrite/promotion

No OCR engine is required: the image model reads visible game text. No actions are silently selected after an HTTP error, malformed response or uncertain view. A visual probe must pass before a real run. The probe checks four-direction image grounding, not Pokémon expertise.

The first implementation calls the VLM on every decision. CLM therefore does not remove visual-perception latency. The emulator pauses during model inference. This is appropriate for experimentation with Pokémon, not a claim of real-time FPS/shooter control. Vision-feature caching and a distilled visual encoder are future optimizations, not measured features of this implementation.

Install and run the already-tested smoke path

On this Mac, the isolated worktree is $WORKSTATION/Projects/mario-rgba-pokemon.

cd "$WORKSTATION/Projects/mario-rgba-pokemon"
bash pixel/setup_mac.sh
pixel/.venv/bin/rgba-game doctor
pixel/.venv/bin/rgba-game smoke --pyboy \
  --out "$WORKSTATION/clm-mac/rgba-smoke-$(date +%Y%m%d-%H%M%S)"

The locked Mac environment uses Python 3.12 and PyBoy 2.7.0. Smoke tests drive a pixel grid through recording, milestone verification, replay and dataset export, then exercise actual PyBoy RGBA capture and save/restore against its bundled test cartridge. These tests do not stand in for Pokémon gameplay results.

Provide a genuine vision model

The old Qwen3-8B CLM backbone is a TEXT embedder. The existing local Qwen3.5-4B RSI quant inspected on this Mac had a vision config but zero vision tensors in its index. Neither is a usable image-reader substitute.

One candidate for a separate image service is the published mlx-community/Qwen3.5-4B-4bit checkpoint through mlx-vlm. The subsequent Pluto experiment loaded this checkpoint, passed the basic image probe, and recorded three failed game trials. It is not a validated Pokémon player; see the measured follow-up. Keep any vision environment separate from MOSS and the existing CLM runtime. Example setup:

python3.12 -m venv "$WORKSTATION/clm-mac/pixel-vision-env"
"$WORKSTATION/clm-mac/pixel-vision-env/bin/python" -m pip install mlx-vlm
"$WORKSTATION/clm-mac/pixel-vision-env/bin/mlx_vlm.server" \
  --host localhost --port 18712 \
  --model mlx-community/Qwen3.5-4B-4bit

Use the full chat-completions URL exposed by that installed server. For an OpenAI-compatible /v1 deployment:

cd "$WORKSTATION/Projects/mario-rgba-pokemon"
export VISION_URL='private workstation service
export VISION_MODEL='mlx-community/Qwen3.5-4B-4bit'
mkdir -p pixel/private
pixel/.venv/bin/rgba-game probe \
  --vision-url "$VISION_URL" --vision-model "$VISION_MODEL" \
  --out pixel/private/vision-proof.json

Keep the service private. Loopback is the default; using a LAN, Tailscale or cloud endpoint requires --allow-remote, which authorizes sending game frames and scene state there. Optional bearer keys come from VLM_API_KEY and CLM_API_KEY, not command-line URL credentials. Models must return the documented JSON contract.

ROM and verifier setup

Use your own local, uncompressed Pokémon Blue .gb file. The adapter checks its header and copies it into the run’s private sandbox; it does not change the original file or an existing cartridge save. Emulator checkpoints use the private run directory. Existing battery saves are not automatically imported.

For reliable milestones, provide a .sym file from the matching ROM build. The profile reads wPartyCount and wObtainedBadges; no guessed addresses or RAM writes are used. Binding a ROM hash records your pairing but does not independently prove that arbitrary supplied symbols match that ROM. Validate the pair before using its results for training.

export BLUE_ROM="$WORKSTATION/Games/Pokemon-Blue.gb"
export BLUE_SYMBOLS="$WORKSTATION/Games/Pokemon-Blue.sym"
pixel/.venv/bin/rgba-game profile \
  --rom "$BLUE_ROM" --symbols "$BLUE_SYMBOLS" \
  --out pixel/private/blue-profile.json

Start with the starter milestone, then move to the Boulder Badge. Do not begin from a state that already satisfies the target: the runtime rejects such trials.

RUN="$WORKSTATION/clm-mac/pokemon-runs/starter-$(date +%Y%m%d-%H%M%S)"
pixel/.venv/bin/rgba-game run \
  --rom "$BLUE_ROM" --out "$RUN" \
  --spec pixel/configs/pokemon_blue_starter.json --goal starter \
  --profile pixel/private/blue-profile.json \
  --vision-url "$VISION_URL" --vision-model "$VISION_MODEL" \
  --vision-proof pixel/private/vision-proof.json \
  --policy teacher --max-steps 200 --max-seconds 900 --quota-mb 256

--window opens PyBoy’s display. --resume /path/to/run/resume.json restores the emulator plus actor memory into a NEW run using the same goal/action spec. For Boulder Badge use pixel/configs/pokemon_blue.json --goal boulder_badge. For eight badges use pixel/configs/pokemon_blue_eight_badges.json --goal eight_badges. Eight badges is explicitly NOT a Hall of Fame or full-game victory. A champion verifier and long-horizon curriculum are not implemented in this version.

Controller timing: directions hold 8 frames and release 8; A/B/Start/Select press 1 frame and release 15; wait advances 32 neutral frames. These are starting settings, not timings already optimized against Blue. Every macro is bounded and opposite-direction combinations are rejected.

Replay, export and train

After a verified teacher milestone, replay its actual recorded controls without calling a model. Every before/after frame and outcome must match:

pixel/.venv/bin/rgba-game replay --run "$RUN" --rom "$BLUE_ROM" \
  --workspace "$RUN-replay-private" --profile pixel/private/blue-profile.json

## Supply successful, independently prepared starting-state groups for one goal.
pixel/.venv/bin/rgba-game export \
  --runs /path/to/teacher-run-A /path/to/teacher-run-B /path/to/teacher-run-C \
  --out pixel/private/blue-teacher.jsonl

Exports include a provenance sidecar. Training re-audits the source evidence and rejects changed or missing proofs. Keep the source run directories alongside the dataset. For transfer between computers, transfer the run artifacts and re-export with local paths; do not hand-edit a provenance file.

Evidence identifiers include measured wall-clock duration, so a semantically identical rerun can produce different byte-level evidence or dataset hashes. Verify each run through its own hash-linked provenance and exact replay rather than expecting rerun SHA-256 values to be identical.

Use the existing CLM/PyTorch environment for training (its dependencies include NumPy, Requests, and Pillow). This task’s validation used an isolated Pillow target at pixel/.test-deps rather than editing the CLM or MOSS environments.

export PYTHONPATH="$PWD/pixel:$PWD/pixel/.test-deps"
"$WORKSTATION/clm-mac/src/CLM/.venv/bin/python" -m rgba_agent train \
  --data pixel/private/blue-teacher.jsonl \
  --checkpoint "$WORKSTATION/.cache/clm/CLM_v0.1-8B.pt" \
  --embed-url private workstation service \
  --encoder-identity 'REPLACE_WITH_ENCODER_WEIGHT_HASH-QUANTIZATION-POOLING-REVISION' \
  --out "$WORKSTATION/clm-mac/pokemon-models/starter-candidate-01" \
  --epochs 100 --device cpu

The embedding endpoint must already serve the matching Qwen3 encoder. It is not a vision endpoint. Give it an immutable, verified identity, not only a mutable model alias. The exact text representation and candidate order use CLM’s existing build_pairs and Embedder implementations. Both projection heads are trained; the original checkpoint is not overwritten. Embedding cache identity includes encoder identity, dataset, scene representation and action descriptions.

Training requires at least THREE distinct starting-state groups, reserves complete groups for model selection, and has no fit-all promotion shortcut. Merely changing a seed label on an otherwise identical start does not create an independent case. Loss or label accuracy never directly promotes a candidate.

Run a student and evaluate it

Serve the newly trained Pokémon head on an unused private port, NOT Mario’s head:

"$WORKSTATION/clm-mac/src/CLM/.venv/bin/clm-serve" \
  --host localhost --port 18711 --device cpu --action-cache 0 \
  --emb-url private workstation service --emb-model qwen3-8b \
  --ckpt "$WORKSTATION/clm-mac/pokemon-models/starter-candidate-01/candidate.pt"

Use run --policy clm --clm-url private workstation service --checkpoint-sha <ACTUAL_SHA256> plus the same ROM, vision and goal arguments. The operator must ensure the endpoint really serves that hash. API aliases alone cannot cryptographically prove remote model identity.

Run candidate and baseline on the same evaluation starts and budgets; replay both. Then rgba-game gate --candidate <run paths...> --baseline <run paths...> --training-report <training-report.json> --out <new gate-report.json> checks exact proofs, stable model/config identities, matched budgets, no training/validation start overlap, and at least TEN distinct scenarios by default. The configurable minimum cannot go below three. A pass is only for that specified scenario suite, not all Pokémon saves or all games. This command writes a report, never replaces the incumbent or automatically starts unlimited training.

Add another game

Implement PixelEnv: observe, step, snapshot, restore, scenario_key, close, and an immutable identity. CallbackRGBA wraps game-specific functions; it does not inject keyboard/mouse events into the operating system. Frame must be an owned HWC uint8 RGBA array. Define a separate action catalogue and success judge. Keep the actor isolated from emulator/object memory.

Games without deterministic save/restore need a different evaluation strategy; they cannot use this exact-replay gate unchanged. Mouse/analog-control games need new bounded actions and an adapter. Hidden information, audio-dependent mechanics, network opponents and strict real-time latency require more than a screenshot interface. Sharing the pipeline does not imply sharing one universal trained policy.

Sources and implementation checks

Local evidence and test counts are in docs/evidence/rgba-pokemon/validation.json.

Continue a verified curriculum without losing the save

Once the starter run is successful and its replay proof passes, explicitly move its checkpoint to the next goal:

pixel/.venv/bin/rgba-game advance --run "$RUN" \
  --next-spec pixel/configs/pokemon_blue.json \
  --out pixel/private/boulder-start.json

Then use a NEW run directory with --goal boulder_badge, the matching spec and --resume pixel/private/boulder-start.json. This retains the actual emulator state and recent memory. It does not label the next milestone complete. The only allowed transitions are starter -> Boulder Badge -> eight badges; action timings must remain identical. Final checkpoints are hash-bound to the run report, so a changed checkpoint invalidates replay/export/transition checks.

Local pixel-head result (2026-09-29)

The local 4.49 MB visual CLM head completed the first-dish task from three nearby starting positions excluded from training and model selection. Full-level and full-game success are not claimed. The separate local Qwen vision experiment passed a basic image probe but stalled in gameplay. See measured results and runnable commands.

The new pixel-feedback perception recipe invalidates old vision-probe identities. Re-run rgba-game probe with the exact selected JSON/scaling/fingerprint options before using the new recipe. Old evidence is preserved, not rewritten.