Decision
The preregistered first strategy failed its public retention floor at all three learning rates. Stop this strategy; do not increase replay coefficients or train through the failed floor. The discarded smoke and all three independent proposals hard-rolled back after their first eight-step evaluation. No checkpoint cleared public selection, so fact confirmation and fresh transfer remain sealed. No fact-confirmation accuracy, research success, strategy success, global promotion or production deployment is claimed.
The implementation and experiment are on branch retention/replay-kd-v1,
stacked on the previously repaired shared-backbone branch. The original tiled
loss file is unchanged. All runs used the same immutable BF16 Qwen3.5-4B
rank-8/alpha-16/dropout-0 parent and existing seven projection modules on one
RTX3090, physical GPU0/logical0, through Desktop Commander on .51.
Actual first-block outcomes, before rollback
| Independent run | Registered peak LR | Steps actually executed | Fact DEV | Regression | Adaptation | Proxy | Decision |
|---|---|---|---|---|---|---|---|
| Parent | — | 0 | 0/24 | 446/446 | 139/240 | 90/144 | Immutable reference |
| Discarded smoke | 2e-5 | 8 | 0/24 | 346/446 | 139/240 | 92/144 | HARD_ROLLBACK |
| Main | 2e-5 | 8 | 0/24 | 346/446 | 139/240 | 92/144 | HARD_ROLLBACK |
| Low | 1e-5 | 8 | 0/24 | 429/446 | 139/240 | 92/144 | HARD_ROLLBACK |
| High | 3e-5 | 8 | 0/24 | 386/446 | 139/240 | 92/144 | HARD_ROLLBACK |
The 96-step schedules were maximum proposal lengths, not a requirement to ignore a safety breach. Three experiment proposals executed 24 optimizer steps in total; the separate excluded smoke added eight. No run reached step16. Each proposal started from the same parent, not the previous proposal’s state. All four exact adapter/AdamW/RNG rollbacks restored step0. The parent was never overwritten. An operational exit0 reports a successfully exercised rejection path, not a successful retention or learning result.
Fact DEV0/24 after eight steps is not a heldout result or a conclusion that the model cannot learn. Retention failed before the registered patience mechanism or later learning checkpoints could be reached. The nonmonotonic7→6→20 test passed; no proposal stopped because fact DEV was flat.
Selective functional failures
| Family | Parent | Main2e-5 | Low1e-5 | High3e-5 |
|---|---|---|---|---|
| RPN | 40/40 | 0/40 | 40/40 | 0/40 |
| LRU | 40/40 | 0/40 | 40/40 | 40/40 |
| BFS | 33/33 | 13/33 | 33/33 | 13/33 |
| Previous smaller | 35/35 | 35/35 | 18/35 | 35/35 |
| Sum-excluding-self | 33/33 | 33/33 | 33/33 | 33/33 |
| Rational proxy | 2/24 | 4/24 | 4/24 | 4/24 |
All other scored public families were unchanged. Main lost100 regression cases; low lost17; high lost60. The proxy’s two recovered rational cases did not offset or excuse any retention loss. Main/high breached RPN sentinel protection, main also breached LRU, and low independently crossed the sub440 hard threshold.
The retained generated programs make the failures inspectable. Main’s RPN
response calls .split() on list input. Its LRU response interprets the
operation list at the wrong nesting level. Low’s previous-smaller response
tracks only the immediately previous item rather than the required predecessor
relation. No evaluator cases or thresholds were modified after seeing these
failures. The same main/smoke plan reproduced the eight component-loss and
gradient records, rejected adapter hash, 31 coding outputs and24 fact outputs
exactly in independent starts.
Teacher and objective verification
The immutable teacher cache contains31 protected sequences and5,235 answer tokens. It stores exact prompt/answer IDs, top64 temperature2 log probabilities, complement mass, weak-margin positions, metadata and teacher/generation hashes. The completed files are read-only and individually hashed. Actual student-equals-teacher KD was zero for all31 sequences. An independent CPU check found at most1.41e-6 probability-normalization error across all top64+tail rows. Teacher tensor data occupy about2.76MB; no full-vocabulary cache lives on GPU.
The coefficients stayed exactly1.0 new CE,1.5 replay CE,0.5 temperature-scaled KD and0.10 margin. Each optimizer step used six new examples and one replay sequence. T2 KL includes the residual bin and T-squared multiplier; replay and new CE each average their own answer tokens. Margins use the registered squared hinge on weak teacher-chosen-versus-competitor margins.
The fixed96-draw replay plan covers all31 families. Fragile sampling is exactly 3x within the ordinary shuffled bag, with every fourth draw additionally forced fragile. All proposals aborted at8, after seeing only seven distinct replay families: range_add, LRU, range, paths, wildcard, CSV and previous-smaller. RPN/BFS were not replayed during that short prefix. This is an observed exposure limitation, not proof that replay ordering alone caused the failures; changing it would require a separate registered experiment.
Per-component gradients are retained rather than hidden behind aggregate loss. For example, main’s mean new-CE gradient norm was14.79; replay CE0.898, KD0.476 and margin3.038 before their respective coefficients. The first KD and margin losses were zero at the parent, as expected. Norms alone do not establish vector conflict, causality or a recommended new coefficient. No coefficient was retuned after inspecting them.
Invariants, isolation and tests
All252 pytest tests passed before training. Coverage includes top64/tail KL, student=teacher near-zero loss, all-token/single-backbone calculations, replay coverage/weights, nonmonotonic fact patience, rollback floors, sentinel damage, exact adapter/optimizer/RNG restoration, one-GPU rejection and exact cold-token comparison. The existing shared-backbone loss file passed its byte-hash lock.
The actual257/700/1025/1536-token numerical fixtures all passed the unchanged 5% relative-gradient and0.02 absolute-loss bounds. Full GPU reference arithmetic was fitted under the11.5GiB allocator cap by diagnostic-only saved-tensor and frozen unused-weight residency offload. A700-token offload bridge reproduced its original GPU reference bit-for-bit. Student arithmetic/loss and BF16 parameter values were not changed. Peak full-reference allocation was11.279GiB; training peaks were about9.216GiB.
Real Landlock workers denied direct and symlink reads of sealed fact and transfer data. Model/cache imports, CUDA and the digest-pinned isolated grader were tested under that restriction. A narrow permission for naming CUDA’s own threads was added after the initial restricted CUDA initialization failed; no Desktop Commander permission was changed. This is direct-read isolation for the reviewed workflow, not a security claim about arbitrary hostile programs with a host Docker socket.
Data order used seed2026093018, independently of the inherited model-runtime initialization seed20260929. Both are recorded, and no extra seed trial occurred. The three proposals used the exact same96-step plan hash and separate fresh parent/optimizer/RNG states. Original model shard and adapter hashes were reverified. PyTorch2.11.0+cu128, CUDA12.8, cuDNN91900 and the full library/device inventory are archived.
Measured work and scope of timing
The discarded smoke measured approximately57.6 answer tokens per training second and204.6 seconds for its complete public evaluation. Teacher cache construction took approximately229.5 seconds. Multiplying that first-block training/evaluation cost by36 blocks gives a conditional2.25-hour model-work projection for three uninterrupted96-step schedules. This excludes reviews, startup, final/cold checks and contention; it is not a measured full run.
Actual main/low/high wall times, including their reviews, were approximately 378.1,397.5 and441.3 seconds because all three stopped at8. There was no need to spend the planned full schedule after explicit hard-floor failures.
Selection boundary and next strategy
select_pareto_candidate.py completed under holdout isolation and returned
NO_PUBLIC_ELIGIBLE_CANDIDATE. There is no final-frozen adapter, and neither the
48 fact confirmation nor144 fresh transfer cases were opened for evaluation.
The prior fact48 recovery check also remains unrun because no selection froze.
These cells are not evaluated, not zero scores and not a failed cold pass.
The preregistered “regression cannot hold at any LR” branch is now active: stop replay-KD v1. Isolated fact LoRA is the next priority rather than a larger replay coefficient. EWC/L2-SP remains a separate proposed comparison requiring its own gradient-norm calibration. None of those alternatives, and no unpinned RKR, was implemented or run here. This result rejects this exact one-adapter recipe and schedule; it is not a universal claim that frozen-teacher KD cannot mitigate forgetting.
Closing verification and parent recovery
The completed smoke and three proposal records were found on the authorised host and independently re-audited before publishing this report. No extra proposal, seed, coefficient change or training-through-rollback was introduced during finalisation. All 252 tests passed again using the existing isolated pytest dependency directory. An initial recheck omitted that directory and failed to import pytest; it ran no tests or model calls and is preserved.
The existing separate cold process restored the safe parent and performed 55 actual generations with zero cache substitutions: 31 coding programs and 24 fact-development answers. Token IDs, response hashes and grades all matched the immutable cache. This reproduced 446/446 regression, 139/240 adaptation, 90/144 proxies and 0/24 new-fact DEV. All four serialized safe checkpoints identify the same restored step-0 parent. This is verified parent recovery, not a cold-tested improved candidate.
Final data/cache/receipt arithmetic audits passed without opening the sealed
sets. The source hash still matches every executed proposal and consumed
approval. Full results, component losses, exact generated programs, rollback
records, public manifests and cold recovery are in
experiments/retention-replay-kd-v1-20260930/. Model binaries, private reasoning,
credentials and the actual sealed confirmation/transfer files are excluded.
The teacher tensor cache remains local, immutable and hash-verified.