Back to the experimentSupporting notebook

Replay plus frozen-parent KD, preregistered retention strategy v1

Source: llm-polisher/retention-replay-kd-v1-20260930/PREREGISTRATION.md · revision 6550ead3945b

Completed result: the smoke and all three registered LRs hard-rolled back at step 8. No candidate qualified and reserved sets remain sealed. See RETENTION_REPLAY_KD_V1_RESULTS.md; the preregistration below is retained unchanged.

This is the user’s first retention experiment, not a new global code-chain promotion. It is stacked on the repaired shared-backbone branch; the existing sequence_tiles.py is byte-hash locked and has not been edited.

Operative parent and objective

The frozen BF16 Qwen3.5-4B backbone and one active rank-8, alpha-16, dropout-0 adapter use the seven existing target projections. The teacher is the same immutable adapter used to start every proposal, never the base model and never the learned-but-forgetting candidate. Its adapter-file SHA-256 is 91c7d7f803eb8be4fc6befbb4cbe28ed0116e3f38274a971d0a46b30f4a90f23. It is known to score 446/446 on the original public coding suite, but was not previously approved as a global successor. Those are separate qualifications.

The objective is fixed before training:

new CE + 1.5 replay CE + 0.5 KD + 0.10 margin.

KD is teacher-to-student KL on 64 cached tokens plus an explicit complement bin at temperature 2, multiplied by temperature squared. It does not discard or renormalize away the tail. CE averages all answer tokens in each sequence. A new microbatch has six facts, a replay microbatch one protected answer. The margin term is a squared hinge against the cached chosen-token versus runner-up margin at positions with teacher margin <=0.5. All these choices are explicit in configs/retention/replay-kd-v1.json, not tuned per proposal.

Cache and replay

Every protected public prompt has exact teacher answer tokens, a prompt/answer hash, top-64 token IDs and log probabilities, tail log probability/mass, weak-margin positions, family ID/fragility, teacher checkpoint hash and registered generation configuration hash. The completed cache is read-only and every metadata/tensor file is independently hashed. No full-vocabulary logits persist on GPU: only transient output-head tiles are constructed. Student=teacher KD is checked on CPU tests and again on every real cached sequence before any optimizer update.

The four fragile families are range, previous-smaller, paths and rational. The ordinary shuffled replay bag has exactly three copies per fragile family versus one per other family. Every fourth draw is additionally forced from a shuffled fragile bag. Thus the 3x claim is conditional on ordinary draws; the forced stratum intentionally increases overall fragile exposure. The fixed 96-draw plan covers all protected families and is identical across LRs. Sentinels LRU, sum and RPN remain sampled.

Incorrect or incomplete teacher outputs on diagnostic families are not labeled correct. Their observed answer tokens are preserved as behavior targets, and an incomplete teacher answer receives no fabricated EOS. A cached output loop is not claimed as a successful program or fact.

Data and isolation

The new Neris registry has 12 fresh entity/value associations. Train is 60 single-service prompts (five phrasings per fact), DEV is 12 singles plus 12 pairs, and confirmation is 24 unseen-wording singles plus 24 unseen pairs. Even reverse-ordered pairs are separated between DEV and confirmation. The facts are taught; the experiment tests retained acquisition and new wording/combinations, not inference of unknown facts.

Six fresh transfer families have 24 inputs each. Their cases and the fact confirmation are registered by hash in sealed/final.json but not committed or supplied to reviewers, training or selection. The old consumed fact-48 is used only as a recovery check after selection/freezing, never as fresh evidence.

Train, cache and selection workers enter unprivileged Landlock restrictions before importing ML code. They cannot read either sealed manifest directly or through an ordinary symlink. The actual restricted process successfully imports the ML stack, initializes CUDA, and uses the pinned Docker grader. CUDA requires permission to name its own threads under /proc/self/task; this narrow resource permission does not expose the sealed directory.

This protects the reviewed workflow against direct accidental reads. It is not a claim to contain arbitrary hostile code with a host Docker daemon. Generated programs still execute in the existing unprivileged, no-network, no-host-mount, digest-pinned grader. No Desktop Commander permission changed.

Proposals and controller

The smoke is eight steps and permanently excluded from selection. A tested retention rollback is an operationally valid smoke outcome, not a retention success; a runtime/gradient/checkpoint failure blocks the full proposals. Then the three independent same-parent proposals run in order: main2e-5, low1e-5, high3e-5. Each has up to96 steps,8 warmup,cosine decay,clip1, and identical data order. No proposal continues another proposal’s weights. Every eight steps measures fact DEV, regression446, adaptation240 and proxy144.

Retention>=444 and proxy drop<=2 permits continuation. Scores440–443 roll back; <440 or sentinel case loss>25% hard-roll back. A rollback ends that independent proposal. Fact DEV may fall or remain flat while retention is safe, for up to four evaluation blocks. This permits 7→6→20 rather than enforcing monotonicity. The coefficients and learning-rate schedules never change within a proposal.

Exact safe snapshots include adapter tensors, AdamW moments/step, learning-rate position/counters, and Python/NumPy/Torch CPU/CUDA RNG states. A rollback verifies full-state identity, not just adapter equality. Locally generated full-state checkpoint files are trusted; downloaded pickles are never loaded.

Only immutable-start, latest-safe, best-new-task, best-retention and public Pareto roles keep materialized adapters, deduplicated by tensor hash. Rejected and dominated candidates retain audit hashes and evaluation records. Public selection is lexicographic: regression, fact DEV, proxy, adaptation, then smaller normalized LoRA displacement. The final selection must score446/446, adaptation

=135 and proxy no worse than parent-2. Its hash freezes before reserved access.

Two success levels, neither a code-chain promotion

Research success: confirmation>=75% and >=8 percentage-point gain, regression 446/446, adaptation>=135, proxy floor maintained, no aggregate fresh-transfer decline, changed adapter and exact cold reproduction. This receives LEARNING_AND_RETENTION_VERIFIED only after all gates.

Strategy success additionally requires at least43/48 facts;46/48 is the target. A38/48 result with446/446 is recorded as retained but suppressed plasticity, not a reason to reduce replay during this proposal. No newly-solved-code-family gate is incorrectly applied to the factual task. No result here advances the global recursive code chain or changes a production service.

11.5 GiB numerical gate

The four original synthetic lengths257,700,1025,1536 retain the5% gradient and 0.02 absolute-loss tolerances. Full reference arithmetic remains on GPU. To fit its diagnostic-only full-vocabulary loss under11.5GiB, unused vision and frozen tied embeddings are temporarily moved to CPU after their forward use; head/loss saved tensors are offloaded without pinned-memory layout conversion. The700-token offload bridge matched the original reference exactly. Student loss, model precision and all parameter values remain unchanged.

All four probes passed; the largest full-reference allocation was11.279GiB. An earlier CPU-CE reference bridge and a pinned-memory offload investigation failed their extra diagnostic checks; those logs are retained, not counted as successful training or silently substituted reference results.

Execution and reporting

Entrypoints are the six requested scripts plus the shared helper. Run pytest, cache build, --check-only, the excluded smoke, then the three proposals. The cache builder and trainer verify parent, configuration, cache and source hashes. Independent Muse maximum-reasoning reviews authorize each exact proposal; review approval never overrides deterministic retention/selection gates.

Record actual smoke tokens per training-second and public-evaluation duration before extrapolating workload. Any extrapolation assumes the remaining blocks cost the same and is not a measured completion time. Libraries, CUDA/driver, physical/logical GPU IDs, seed, per-component losses/gradient norms, per-layer LoRA displacement and termination reasons are retained with the run evidence. EWC/L2-SP, isolated adapters and unpinned RKR remain unimplemented.