A smaller speech model earns its place on the iPhone
Aidana transcribed the 47 matched clips 3.35 times faster, with observed English WER of 2.15% and Spanish WER of 5.37%.
4 min readWhat I tried. What happened. What I’d do next.
Showing all 44 experiments
Aidana transcribed the 47 matched clips 3.35 times faster, with observed English WER of 2.15% and Spanish WER of 5.37%.
4 min readThree live recovery scenarios passed, each finishing with 10/10 checks.
4 min readOn Ubuntu CPU, the corrected policy passed 10/10 full-kitchen trials against a same-host retraining control at 7/10.
6 min readAll three registered learning rates crossed the retention floor after eight steps and rolled back.
7 min readThe corrected candidate passed 10/10 paired two-dish trials against a frozen baseline at 0/10.
6 min readThe v3 protocol measured 8/10 full-stage clears versus 0/10 for its paired baseline; the strict gate rejected it.
6 min readThe first-dish baseline produced six successful recorded evaluations across five starting positions, with exact replay.
4 min readOwned frames, bounded controls, saved state, teacher-data export and exact replay form one auditable pipeline.
9 min readThe evaluator tightened survival, terminal-state and exact-replay checks without rewriting historical experimental outcomes.
4 min readBonsai-2-27B PQ2_0 with MTP passed 48/50 tasks with thinking on, at a 15.8-second median on an RTX 3060.
10 min readThe classifier matched its teacher on 192/202 examples, with macro-F1 0.863.
7 min readA real masked object removal finished in 24.39 seconds and preserved every unselected decoded pixel exactly.
3 min readThree initial adapters completed: 600 LTX steps and two Qwen identity runs of 800 steps each.
4 min read213 files became 97 unique decoded images; the curated random selection used 15 distinct photographs or photo-like images.
3 min readA keyboard-specific handler preserved the lock gesture and the display reached OFF roughly 0.7 seconds later.
4 min read25 distinct images were cleaned into lossless copies, preserving dimensions and all decoded pixels outside each mask.
3 min readDecode dropped from 3.36 seconds to 0.24; total render wall time dropped from 35.1 to 24.4 seconds.
2 min readThe binary screen accepted 24/25 cleaned images; both known positive controls were detected.
3 min readEight matched 1024-square scenes averaged 93.352 seconds with FP8 and 46.219 with opt-in INT8.
15 min readA 551-character, eight-sentence input rendered in two chunks with 48.0 seconds of healthy audio.
2 min readThe new 1.76GB FP8-E4M3 codec reached FP32-control parity in the measured round-trip test with its encoder intact.
2 min readUsing FP32 reference encoding lifted AWQ WavLM target similarity from 0.8475 to 0.9556 while retaining INT8 output decode.
3 min readThe balanced two-GPU path took 48.1 seconds but every image was invalid flat gray noise.
2 min readGenerator precision affected the measured ranking; changing the output codec shifted similarity by at most about 0.001.
3 min readAn explicit 12/12/12 GPU expert split plus 12 CPU blocks measured 31.1 tokens per second.
3 min readThe patched vLLM stack measured 67.4 tokens per second for one stream and 73.1 aggregate for two.
3 min readThe recorded service used about 10.3GB resident memory; first render took 14.8 seconds and warm render 5.7.
2 min readThe full-spec job stopped after six of 36 blocks; the monitor detected the dead wrapper and idle GPU.
2 min readAWQ passed 33/33 prompts with pooled real-time factor 0.43 and 35.8 tokens per second.
2 min readThe 1B encoder embedded 1,004 documents per second; the 8B encoder reached 181, with no proven retrieval-quality win.
12 min readPure TTS legs shared token sequences, and normalized transcription differences were zero across the tested codec legs.
3 min readThe HQ4 preview matched 96.5% of text tokens in the greedy test and passed three staged-memory probes at 6.7–6.8 GiB.
3 min readThe staged run peaked at 6.63 GiB, down from 13.01 before staging; 252 INT4 projections passed the verifier.
4 min readDraft length three reached a short-prose mean of 88.16 tokens per second, versus 64.67 with drafting off.
8 min readThe deployment notebook records the EXL3 serving family, cache choices and real runtime constraints.
7 min readThe notebook follows official BF16 weights through EXL3 conversion, artifact checks and deployment work.
4 min readThe promoted merge passed 107 tests with one environmental skip.
8 min readPolish step 400 measured 6.1% word error rate; the final step 1500 measured 42.9%.
2 min readThe clean agentic importance matrix recovered the hard T3 task; later draft decoding improved measured throughput.
7 min readThe rollback fix brought draft decoding to parity at temperature 0.7 and a 24% gain at greedy sampling.
7 min readThe warm proxy decode median was 59.17 tokens per second; a 198,960-token request passed with GPU headroom.
4 min readOrnith measured 210.16 tokens per second on the realistic 256-token prompt and 270.40 on predictable prose.
5 min readSharp reduced completion tokens by 20.6%, 20.1% and 16.0% across the three deployments.
9 min readAt 960×544, dual INT8 completed a 124-frame, 20-step run in 354.580 seconds.
5 min readTry another term or reset the filters to open the complete notebook.