Completed result: the smoke and all three registered LRs hard-rolled back at step 8. No candidate qualified and reserved sets remain sealed. See
RETENTION_REPLAY_KD_V1_RESULTS.md; the preregistration below is retained unchanged.
This is the user’s first retention experiment, not a new global code-chain
promotion. It is stacked on the repaired shared-backbone branch; the existing
sequence_tiles.py is byte-hash locked and has not been edited.
Operative parent and objective
The frozen BF16 Qwen3.5-4B backbone and one active rank-8, alpha-16, dropout-0
adapter use the seven existing target projections. The teacher is the same
immutable adapter used to start every proposal, never the base model and never
the learned-but-forgetting candidate. Its adapter-file SHA-256 is
91c7d7f803eb8be4fc6befbb4cbe28ed0116e3f38274a971d0a46b30f4a90f23.
It is known to score 446/446 on the original public coding suite, but was not
previously approved as a global successor. Those are separate qualifications.
The objective is fixed before training:
new CE + 1.5 replay CE + 0.5 KD + 0.10 margin.
KD is teacher-to-student KL on 64 cached tokens plus an explicit complement
bin at temperature 2, multiplied by temperature squared. It does not discard
or renormalize away the tail. CE averages all answer tokens in each sequence.
A new microbatch has six facts, a replay microbatch one protected answer.
The margin term is a squared hinge against the cached chosen-token versus
runner-up margin at positions with teacher margin <=0.5. All these choices
are explicit in configs/retention/replay-kd-v1.json, not tuned per proposal.
Cache and replay
Every protected public prompt has exact teacher answer tokens, a prompt/answer hash, top-64 token IDs and log probabilities, tail log probability/mass, weak-margin positions, family ID/fragility, teacher checkpoint hash and registered generation configuration hash. The completed cache is read-only and every metadata/tensor file is independently hashed. No full-vocabulary logits persist on GPU: only transient output-head tiles are constructed. Student=teacher KD is checked on CPU tests and again on every real cached sequence before any optimizer update.
The four fragile families are range, previous-smaller, paths and rational. The ordinary shuffled replay bag has exactly three copies per fragile family versus one per other family. Every fourth draw is additionally forced from a shuffled fragile bag. Thus the 3x claim is conditional on ordinary draws; the forced stratum intentionally increases overall fragile exposure. The fixed 96-draw plan covers all protected families and is identical across LRs. Sentinels LRU, sum and RPN remain sampled.
Incorrect or incomplete teacher outputs on diagnostic families are not labeled correct. Their observed answer tokens are preserved as behavior targets, and an incomplete teacher answer receives no fabricated EOS. A cached output loop is not claimed as a successful program or fact.
Data and isolation
The new Neris registry has 12 fresh entity/value associations. Train is 60 single-service prompts (five phrasings per fact), DEV is 12 singles plus 12 pairs, and confirmation is 24 unseen-wording singles plus 24 unseen pairs. Even reverse-ordered pairs are separated between DEV and confirmation. The facts are taught; the experiment tests retained acquisition and new wording/combinations, not inference of unknown facts.
Six fresh transfer families have 24 inputs each. Their cases and the fact
confirmation are registered by hash in sealed/final.json but not committed
or supplied to reviewers, training or selection. The old consumed fact-48 is
used only as a recovery check after selection/freezing, never as fresh evidence.
Train, cache and selection workers enter unprivileged Landlock restrictions
before importing ML code. They cannot read either sealed manifest directly
or through an ordinary symlink. The actual restricted process successfully
imports the ML stack, initializes CUDA, and uses the pinned Docker grader.
CUDA requires permission to name its own threads under /proc/self/task;
this narrow resource permission does not expose the sealed directory.
This protects the reviewed workflow against direct accidental reads. It is not a claim to contain arbitrary hostile code with a host Docker daemon. Generated programs still execute in the existing unprivileged, no-network, no-host-mount, digest-pinned grader. No Desktop Commander permission changed.
Proposals and controller
The smoke is eight steps and permanently excluded from selection. A tested retention rollback is an operationally valid smoke outcome, not a retention success; a runtime/gradient/checkpoint failure blocks the full proposals. Then the three independent same-parent proposals run in order: main2e-5, low1e-5, high3e-5. Each has up to96 steps,8 warmup,cosine decay,clip1, and identical data order. No proposal continues another proposal’s weights. Every eight steps measures fact DEV, regression446, adaptation240 and proxy144.
Retention>=444 and proxy drop<=2 permits continuation. Scores440–443 roll back; <440 or sentinel case loss>25% hard-roll back. A rollback ends that independent proposal. Fact DEV may fall or remain flat while retention is safe, for up to four evaluation blocks. This permits 7→6→20 rather than enforcing monotonicity. The coefficients and learning-rate schedules never change within a proposal.
Exact safe snapshots include adapter tensors, AdamW moments/step, learning-rate position/counters, and Python/NumPy/Torch CPU/CUDA RNG states. A rollback verifies full-state identity, not just adapter equality. Locally generated full-state checkpoint files are trusted; downloaded pickles are never loaded.
Only immutable-start, latest-safe, best-new-task, best-retention and public Pareto roles keep materialized adapters, deduplicated by tensor hash. Rejected and dominated candidates retain audit hashes and evaluation records. Public selection is lexicographic: regression, fact DEV, proxy, adaptation, then smaller normalized LoRA displacement. The final selection must score446/446, adaptation
=135 and proxy no worse than parent-2. Its hash freezes before reserved access.
Two success levels, neither a code-chain promotion
Research success: confirmation>=75% and >=8 percentage-point gain, regression
446/446, adaptation>=135, proxy floor maintained, no aggregate fresh-transfer
decline, changed adapter and exact cold reproduction. This receives
LEARNING_AND_RETENTION_VERIFIED only after all gates.
Strategy success additionally requires at least43/48 facts;46/48 is the target. A38/48 result with446/446 is recorded as retained but suppressed plasticity, not a reason to reduce replay during this proposal. No newly-solved-code-family gate is incorrectly applied to the factual task. No result here advances the global recursive code chain or changes a production service.
11.5 GiB numerical gate
The four original synthetic lengths257,700,1025,1536 retain the5% gradient and 0.02 absolute-loss tolerances. Full reference arithmetic remains on GPU. To fit its diagnostic-only full-vocabulary loss under11.5GiB, unused vision and frozen tied embeddings are temporarily moved to CPU after their forward use; head/loss saved tensors are offloaded without pinned-memory layout conversion. The700-token offload bridge matched the original reference exactly. Student loss, model precision and all parameter values remain unchanged.
All four probes passed; the largest full-reference allocation was11.279GiB. An earlier CPU-CE reference bridge and a pinned-memory offload investigation failed their extra diagnostic checks; those logs are retained, not counted as successful training or silently substituted reference results.
Execution and reporting
Entrypoints are the six requested scripts plus the shared helper. Run pytest,
cache build, --check-only, the excluded smoke, then the three proposals. The
cache builder and trainer verify parent, configuration, cache and source hashes.
Independent Muse maximum-reasoning reviews authorize each exact proposal;
review approval never overrides deterministic retention/selection gates.
Record actual smoke tokens per training-second and public-evaluation duration before extrapolating workload. Any extrapolation assumes the remaining blocks cost the same and is not a measured completion time. Libraries, CUDA/driver, physical/logical GPU IDs, seed, per-component losses/gradient norms, per-layer LoRA displacement and termination reasons are retained with the run evidence. EWC/L2-SP, isolated adapters and unpinned RKR remain unimplemented.