The experiment archive

The journal.

What I tried. What happened. What I’d do next.

Showing all 44 experiments

04Training & adaptationRejected

When the retention gate says stop

All three registered learning rates crossed the retention floor after eight steps and rolled back.

7 min read
05Game learningMeasured

A second dish, a new frozen test

The corrected candidate passed 10/10 paired two-dish trials against a frozen baseline at 0/10.

6 min read
06Game learningRejected

Eight clears are still not a promotion

The v3 protocol measured 8/10 full-stage clears versus 0/10 for its paired baseline; the strict gate rejected it.

6 min read
07Game learningMeasured

Learning to play from pixels

The first-dish baseline produced six successful recorded evaluations across five starting positions, with exact replay.

4 min read
08Game learningField notes

The reusable RGBA learning pipeline

Owned frames, bounded controls, saved state, teacher-data export and exact replay form one auditable pipeline.

9 min read
09Game learningField notes

What counts as a successful episode?

The evaluator tightened survival, terminal-state and exact-replay checks without rewriting historical experimental outcomes.

4 min read
10Language modelsMeasured

Finding a local model for the inbox

Bonsai-2-27B PQ2_0 with MTP passed 48/50 tasks with thinking on, at a 15.8-second median on an RTX 3060.

10 min read
16Images & videoMeasured

Masked cleanup before a fresh LoRA

25 distinct images were cleaned into lossless copies, preserving dimensions and all decoded pixels outside each mask.

3 min read
17Speech & codecsMeasured

Moving speech decode onto the GPU

Decode dropped from 3.36 seconds to 0.24; total render wall time dropped from 35.1 to 24.4 seconds.

2 min read
20Speech & codecsMeasured

Long speech, bounded memory

A 551-character, eight-sentence input rendered in two chunks with 48.0 seconds of healthy audio.

2 min read
21Speech & codecsMeasured

Round-trip tests expose a broken encoder

The new 1.76GB FP8-E4M3 codec reached FP32-control parity in the measured round-trip test with its encoder intact.

2 min read
22Speech & codecsMeasured

The clone bug was in the reference encoder

Using FP32 reference encoding lifted AWQ WavLM target similarity from 0.8475 to 0.9556 while retaining INT8 output decode.

3 min read
28Speech & codecsIncident

A quantization job goes silent

The full-spec job stopped after six of 36 blocks; the monitor detected the dead wrapper and idle GPU.

2 min read
32Speech & codecsMeasured

Keeping the audio-sensitive layers intact

The HQ4 preview matched 96.5% of text tokens in the greedy test and passed three staged-memory probes at 6.7–6.8 GiB.

3 min read
34Language modelsMeasured

Three draft tokens hit the sweet spot

Draft length three reached a short-prose mean of 88.16 tokens per second, versus 64.67 with drafting off.

8 min read
36Language modelsIn progress

The EXL3 quantization journey

The notebook follows official BF16 weights through EXL3 conversion, artifact checks and deployment work.

4 min read
41Language modelsMeasured

Making room for a 200K context

The warm proxy decode median was 59.17 tokens per second; a 198,960-token request passed with GPU headroom.

4 min read
42Language modelsMeasured

DFlash2 on a single RTX 3090

Ornith measured 210.16 tokens per second on the realistic 256-token prompt and 270.40 on predictable prose.

5 min read