Figures from the workbench The results. The measurements, with the conditions still attached.
Every figure comes from a recorded measurement or an explicit report table. Compare within the stated workload, inspect the limits, and take the data with you.
How I report results Research area Every area Agents & systems Game learning Images & video Language models Speech & codecs Training & adaptation
Showing 34 figures
Recorded result What recovery cost in three live scenarios Single RTX 3090 · local Qwen worker + GLM reviewer · each scenario finished 10/10 checks
What recovery cost in three live scenarios Repeated failures · 2 reviews: 150 seconds; Acceptance failure · 1 review: 142 seconds; Successful task · no review: 9 seconds. Deliberately triggered integration scenarios; duration does not measure general coding accuracy. Repeated failures · 2 reviews 150 Repeated failures · 2 reviews: 150 seconds Acceptance failure · 1 review 142 Acceptance failure · 1 review: 142 seconds Successful task · no review 9 Successful task · no review: 9 seconds 0 seconds
Repeated failures · 2 reviews 150
Acceptance failure · 1 review 142
Successful task · no review 9
seconds
Deliberately triggered integration scenarios; duration does not measure general coding accuracy.
View data & source What recovery cost in three live scenarios · seconds Configuration Value Repeated failures · 2 reviews 150 Acceptance failure · 1 review 142 Successful task · no review 9
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result A compact handoff can still add information Serialized context at the recovery handoff
A compact handoff can still add information Repeated failures · before: 14,325 characters; Repeated failures · after: 5,723 characters; Acceptance failure · before: 5,290 characters; Acceptance failure · after: 6,566 characters. The short acceptance request grew after adding diagnostic information; this is not a universal compression ratio. Repeated failures · before 14,325 Repeated failures · before: 14,325 characters Repeated failures · after 5,723 Repeated failures · after: 5,723 characters Acceptance failure · before 5,290 Acceptance failure · before: 5,290 characters Acceptance failure · after 6,566 Acceptance failure · after: 6,566 characters 0 characters
Repeated failures · before 14,325
Repeated failures · after 5,723
Acceptance failure · before 5,290
Acceptance failure · after 6,566
characters
The short acceptance request grew after adding diagnostic information; this is not a universal compression ratio.
View data & source A compact handoff can still add information · characters Configuration Value Repeated failures · before 14,325 Repeated failures · after 5,723 Acceptance failure · before 5,290 Acceptance failure · after 6,566
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Every candidate crossed the retention floor Public regression set · first evaluation block · no confirmation/transfer reads
Every candidate crossed the retention floor Frozen parent: 446 checks passed / 446; LR 1e-5 · after 8 steps: 429 checks passed / 446; LR 2e-5 · after 8 steps: 346 checks passed / 446; LR 3e-5 · after 8 steps: 386 checks passed / 446; Restored parent: 446 checks passed / 446. All three runs rolled back. A restored parent is not an improved trained checkpoint. Frozen parent 446 Frozen parent: 446 checks passed / 446 LR 1e-5 · after 8 steps 429 LR 1e-5 · after 8 steps: 429 checks passed / 446 LR 2e-5 · after 8 steps 346 LR 2e-5 · after 8 steps: 346 checks passed / 446 LR 3e-5 · after 8 steps 386 LR 3e-5 · after 8 steps: 386 checks passed / 446 Restored parent 446 Restored parent: 446 checks passed / 446 Gate: 440 0 checks passed / 446
Frozen parent 446
LR 1e-5 · after 8 steps 429
LR 2e-5 · after 8 steps 346
LR 3e-5 · after 8 steps 386
Restored parent 446
checks passed / 446 · gate: 440
All three runs rolled back. A restored parent is not an improved trained checkpoint.
View data & source Every candidate crossed the retention floor · checks passed / 446 Configuration Value Frozen parent 446 LR 1e-5 · after 8 steps 429 LR 2e-5 · after 8 steps 346 LR 3e-5 · after 8 steps 386 Restored parent 446
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The smaller encoder moved more documents 275 documents · batch size 32 · AWQ W4A16
The smaller encoder moved more documents 1B · native 2048d: 1,003.8 documents / second; 8B · native 4096d: 181.0 documents / second. Amortized throughput on this corpus, not single-query latency. 1B · native 2048d 1,003.8 1B · native 2048d: 1,003.8 documents / second 8B · native 4096d 181.0 8B · native 4096d: 181.0 documents / second 0 documents / second
1B · native 2048d 1,003.8
8B · native 4096d 181.0
documents / second
Amortized throughput on this corpus, not single-query latency.
View data & source The smaller encoder moved more documents · documents / second Configuration Value 1B · native 2048d 1,003.8 8B · native 4096d 181.0
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Memory above the idle GPU baseline Peak GPU allocation delta · 4 MiB baseline
Memory above the idle GPU baseline 1B: 2,380 MiB; 8B: 7,530 MiB. Native dimensions differ. Index storage is a separate cost. 1B 2,380 1B: 2,380 MiB 8B 7,530 8B: 7,530 MiB 0 MiB
Native dimensions differ. Index storage is a separate cost.
View data & source Memory above the idle GPU baseline · MiB Configuration Value 1B 2,380 8B 7,530
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Retrieval scores were close on the bilingual set 110 paired queries · 275 documents · English and Spanish
Retrieval scores were close on the bilingual set 1B · native 2048d: 0.955 MRR@10; 8B · native 4096d: 0.962 MRR@10; 8B · sliced 2048d: 0.971 MRR@10. Neither paired bootstrap interval establishes a quality win; see the difference plot. 1B · native 2048d 0.955 1B · native 2048d: 0.955 MRR@10 8B · native 4096d 0.962 8B · native 4096d: 0.962 MRR@10 8B · sliced 2048d 0.971 8B · sliced 2048d: 0.971 MRR@10 0 MRR@10
1B · native 2048d 0.955
8B · native 4096d 0.962
8B · sliced 2048d 0.971
MRR@10
Neither paired bootstrap interval establishes a quality win; see the difference plot.
View data & source Retrieval scores were close on the bilingual set · MRR@10 Configuration Value 1B · native 2048d 0.955 8B · native 4096d 0.962 8B · sliced 2048d 0.971
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Both quality intervals cross zero 8B − 1B · paired 95% bootstrap intervals · 110 queries
Both quality intervals cross zero 8B · sliced_2048d vs 1B: 0.0167 MRR@10 difference, 95 percent interval -0.0061 to 0.0432; 8B · native_4096d vs 1B: 0.0076 MRR@10 difference, 95 percent interval -0.0189 to 0.0371. An interval crossing zero does not establish equivalence or a positive quality gain. 8B · sliced_2048d vs 1B 0.0167 95% interval: -0.0061 to 0.0432 8B · native_4096d vs 1B 0.0076 95% interval: -0.0189 to 0.0371 Zero difference
8B · sliced_2048d vs 1B 0.0167
95% interval: -0.0061 to 0.0432 · zero marked
8B · native_4096d vs 1B 0.0076
95% interval: -0.0189 to 0.0371 · zero marked
An interval crossing zero does not establish equivalence or a positive quality gain.
View data & source Both quality intervals cross zero · MRR@10 difference Configuration Value 95% low 95% high 8B · sliced_2048d vs 1B 0.0167 -0.0061 0.0432 8B · native_4096d vs 1B 0.0076 -0.0189 0.0371
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Three draft tokens hit the sweet spot Short-prose mean · EXL3 · Q8 cache · three RTX 3090s
Three draft tokens hit the sweet spot Drafting off: 64.7 tokens / second; 2 draft tokens: 84.4 tokens / second; 3 draft tokens: 88.2 tokens / second; 4 draft tokens: 80.7 tokens / second; 6 draft tokens: 62.4 tokens / second. The separate longer prompt and concurrent streams are different workloads. Drafting off 64.7 Drafting off: 64.7 tokens / second 2 draft tokens 84.4 2 draft tokens: 84.4 tokens / second 3 draft tokens 88.2 3 draft tokens: 88.2 tokens / second 4 draft tokens 80.7 4 draft tokens: 80.7 tokens / second 6 draft tokens 62.4 6 draft tokens: 62.4 tokens / second 0 tokens / second
Drafting off 64.7
2 draft tokens 84.4
3 draft tokens 88.2
4 draft tokens 80.7
6 draft tokens 62.4
tokens / second
The separate longer prompt and concurrent streams are different workloads.
View data & source Three draft tokens hit the sweet spot · tokens / second Configuration Value Drafting off 64.7 2 draft tokens 84.4 3 draft tokens 88.2 4 draft tokens 80.7 6 draft tokens 62.4
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The best checkpoint was not the last one Five fixed Spanish prompts · 49 reference words · STT screen
The best checkpoint was not the last one Step 100: 10.2 word error rate (%); Step 200: 14.3 word error rate (%); Step 300: 26.5 word error rate (%); Step 400: 6.1 word error rate (%); Step 500: 42.9 word error rate (%); Step 600: 26.5 word error rate (%); Step 700: 18.4 word error rate (%); Step 800: 24.5 word error rate (%); Step 900: 12.2 word error rate (%); Step 1000: 16.3 word error rate (%); Step 1100: 14.3 word error rate (%); Step 1200: 16.3 word error rate (%); Step 1300: 28.6 word error rate (%); Step 1400: 28.6 word error rate (%); Step 1500: 42.9 word error rate (%). Transcript error does not measure accent, timbre or listener preference. 0% 12% 24% 36% 48% Step 100: 10.2% Step 200: 14.3% Step 300: 26.5% Step 400: 6.1% Best: 6.1% Step 500: 42.9% Step 600: 26.5% Step 700: 18.4% Step 800: 24.5% Step 900: 12.2% Step 1000: 16.3% Step 1100: 14.3% Step 1200: 16.3% Step 1300: 28.6% Step 1400: 28.6% Step 1500: 42.9% 100 500 1000 1500 Training step · word error rate (%) 0% 24% 48% Step 100: 10.2% Step 200: 14.3% Step 300: 26.5% Step 400: 6.1% Step 500: 42.9% Step 600: 26.5% Step 700: 18.4% Step 800: 24.5% Step 900: 12.2% Step 1000: 16.3% Step 1100: 14.3% Step 1200: 16.3% Step 1300: 28.6% Step 1400: 28.6% Step 1500: 42.9% 100 500 1000 1500 Training step Transcript error does not measure accent, timbre or listener preference.
View data & source The best checkpoint was not the last one · word error rate (%) Configuration Value Step 100 10.2 Step 200 14.3 Step 300 26.5 Step 400 6.1 Step 500 42.9 Step 600 26.5 Step 700 18.4 Step 800 24.5 Step 900 12.2 Step 1000 16.3 Step 1100 14.3 Step 1200 16.3 Step 1300 28.6 Step 1400 28.6 Step 1500 42.9
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Matched scenes, different diffusion precision 1024×1024 · 40 steps · matched seeds 2401–2408 · warm RTX 3090 runs
Matched scenes, different diffusion precision FP8 · mean of 8 scenes: 93.352 seconds; INT8 · mean of 8 scenes: 46.219 seconds. INT8 was opt-in. Fine detail changed, so timing improvement is not visual-quality parity. FP8 · mean of 8 scenes 93.352 FP8 · mean of 8 scenes: 93.352 seconds INT8 · mean of 8 scenes 46.219 INT8 · mean of 8 scenes: 46.219 seconds 0 seconds
FP8 · mean of 8 scenes 93.352
INT8 · mean of 8 scenes 46.219
seconds
INT8 was opt-in. Fine detail changed, so timing improvement is not visual-quality parity.
View data & source Matched scenes, different diffusion precision · seconds Configuration Value FP8 · mean of 8 scenes 93.352 INT8 · mean of 8 scenes 46.219
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Every matched scene in the timing suite Same eight scenes · FP8 and INT8 · 40 steps
Every matched scene in the timing suite Seed 2401 · FP8: 93.215 seconds; Seed 2401 · INT8: 46.079 seconds; Seed 2402 · FP8: 93.244 seconds; Seed 2402 · INT8: 46.154 seconds; Seed 2403 · FP8: 93.446 seconds; Seed 2403 · INT8: 46.262 seconds; Seed 2404 · FP8: 93.395 seconds; Seed 2404 · INT8: 46.265 seconds; Seed 2405 · FP8: 93.318 seconds; Seed 2405 · INT8: 46.217 seconds; Seed 2406 · FP8: 93.357 seconds; Seed 2406 · INT8: 46.271 seconds; Seed 2407 · FP8: 93.449 seconds; Seed 2407 · INT8: 46.274 seconds; Seed 2408 · FP8: 93.389 seconds; Seed 2408 · INT8: 46.227 seconds. Each pair shares a seed and workload. Error bars are unavailable; these are recorded runs. Seed 2401 · FP8 93.215 Seed 2401 · FP8: 93.215 seconds Seed 2401 · INT8 46.079 Seed 2401 · INT8: 46.079 seconds Seed 2402 · FP8 93.244 Seed 2402 · FP8: 93.244 seconds Seed 2402 · INT8 46.154 Seed 2402 · INT8: 46.154 seconds Seed 2403 · FP8 93.446 Seed 2403 · FP8: 93.446 seconds Seed 2403 · INT8 46.262 Seed 2403 · INT8: 46.262 seconds Seed 2404 · FP8 93.395 Seed 2404 · FP8: 93.395 seconds Seed 2404 · INT8 46.265 Seed 2404 · INT8: 46.265 seconds Seed 2405 · FP8 93.318 Seed 2405 · FP8: 93.318 seconds Seed 2405 · INT8 46.217 Seed 2405 · INT8: 46.217 seconds Seed 2406 · FP8 93.357 Seed 2406 · FP8: 93.357 seconds Seed 2406 · INT8 46.271 Seed 2406 · INT8: 46.271 seconds Seed 2407 · FP8 93.449 Seed 2407 · FP8: 93.449 seconds Seed 2407 · INT8 46.274 Seed 2407 · INT8: 46.274 seconds Seed 2408 · FP8 93.389 Seed 2408 · FP8: 93.389 seconds Seed 2408 · INT8 46.227 Seed 2408 · INT8: 46.227 seconds 0 seconds
Seed 2401 · FP8 93.215
Seed 2401 · INT8 46.079
Seed 2402 · FP8 93.244
Seed 2402 · INT8 46.154
Seed 2403 · FP8 93.446
Seed 2403 · INT8 46.262
Seed 2404 · FP8 93.395
Seed 2404 · INT8 46.265
Seed 2405 · FP8 93.318
Seed 2405 · INT8 46.217
Seed 2406 · FP8 93.357
Seed 2406 · INT8 46.271
Seed 2407 · FP8 93.449
Seed 2407 · INT8 46.274
Seed 2408 · FP8 93.389
Seed 2408 · INT8 46.227
seconds
Each pair shares a seed and workload. Error bars are unavailable; these are recorded runs.
View data & source Every matched scene in the timing suite · seconds Configuration Value Seed 2401 · FP8 93.215 Seed 2401 · INT8 46.079 Seed 2402 · FP8 93.244 Seed 2402 · INT8 46.154 Seed 2403 · FP8 93.446 Seed 2403 · INT8 46.262 Seed 2404 · FP8 93.395 Seed 2404 · INT8 46.265 Seed 2405 · FP8 93.318 Seed 2405 · INT8 46.217 Seed 2406 · FP8 93.357 Seed 2406 · INT8 46.271 Seed 2407 · FP8 93.449 Seed 2407 · INT8 46.274 Seed 2408 · FP8 93.389 Seed 2408 · INT8 46.227
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result A five-second video at 960×544 Three RTX 3090s · 124 frames · 24 fps · 20 steps · full-decode checks passed
A five-second video at 960×544 W4A8 / NVFP4 · REF2VA: 319.045 elapsed seconds; W4A8 / INT8 · multishot: 321.739 elapsed seconds; W4A8 / NVFP4 · REF2VA · alt. placement: 331.366 elapsed seconds; INT8 / INT8 · T2VA/FL2VA: 354.580 elapsed seconds; INT4Q / INT8 · multishot: 380.578 elapsed seconds; INT8 / INT8 · REF2VA · HQ: 399.403 elapsed seconds; BF16 / Q4_K_M · REF2VA: 501.095 elapsed seconds. Observed production runs: mode, placement and measurement method vary. Read each profile before choosing. W4A8 / NVFP4 · REF2VA 319.045 W4A8 / NVFP4 · REF2VA: 319.045 elapsed seconds W4A8 / INT8 · multishot 321.739 W4A8 / INT8 · multishot: 321.739 elapsed seconds W4A8 / NVFP4 · REF2VA · alt. placement 331.366 W4A8 / NVFP4 · REF2VA · alt. placement: 331.366 elapsed seconds INT8 / INT8 · T2VA/FL2VA 354.580 INT8 / INT8 · T2VA/FL2VA: 354.580 elapsed seconds INT4Q / INT8 · multishot 380.578 INT4Q / INT8 · multishot: 380.578 elapsed seconds INT8 / INT8 · REF2VA · HQ 399.403 INT8 / INT8 · REF2VA · HQ: 399.403 elapsed seconds BF16 / Q4_K_M · REF2VA 501.095 BF16 / Q4_K_M · REF2VA: 501.095 elapsed seconds 0 elapsed seconds
W4A8 / NVFP4 · REF2VA 319.045
W4A8 / INT8 · multishot 321.739
W4A8 / NVFP4 · REF2VA · alt. placement 331.366
INT8 / INT8 · T2VA/FL2VA 354.580
INT4Q / INT8 · multishot 380.578
INT8 / INT8 · REF2VA · HQ 399.403
BF16 / Q4_K_M · REF2VA 501.095
elapsed seconds
Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.
View data & source A five-second video at 960×544 · elapsed seconds Configuration Value W4A8 / NVFP4 · REF2VA 319.045 W4A8 / INT8 · multishot 321.739 W4A8 / NVFP4 · REF2VA · alt. placement 331.366 INT8 / INT8 · T2VA/FL2VA 354.580 INT4Q / INT8 · multishot 380.578 INT8 / INT8 · REF2VA · HQ 399.403 BF16 / Q4_K_M · REF2VA 501.095
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result A five-second video at 1344×768 Three RTX 3090s · 124 frames · 24 fps · 20 steps · full-decode checks passed
A five-second video at 1344×768 W4A8 / W4A8 · REF2VA: 970.306 elapsed seconds; INT8 / INT8 · T2VA/FL2VA: 990.640 elapsed seconds; W4A8 / INT8 · REF2VA: 1,017.780 elapsed seconds; INT8 / INT8 · REF2VA: 1,042.150 elapsed seconds. Observed production runs: mode, placement and measurement method vary. Read each profile before choosing. W4A8 / W4A8 · REF2VA 970.306 W4A8 / W4A8 · REF2VA: 970.306 elapsed seconds INT8 / INT8 · T2VA/FL2VA 990.640 INT8 / INT8 · T2VA/FL2VA: 990.640 elapsed seconds W4A8 / INT8 · REF2VA 1,017.780 W4A8 / INT8 · REF2VA: 1,017.780 elapsed seconds INT8 / INT8 · REF2VA 1,042.150 INT8 / INT8 · REF2VA: 1,042.150 elapsed seconds 0 elapsed seconds
W4A8 / W4A8 · REF2VA 970.306
INT8 / INT8 · T2VA/FL2VA 990.640
W4A8 / INT8 · REF2VA 1,017.780
INT8 / INT8 · REF2VA 1,042.150
elapsed seconds
Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.
View data & source A five-second video at 1344×768 · elapsed seconds Configuration Value W4A8 / W4A8 · REF2VA 970.306 INT8 / INT8 · T2VA/FL2VA 990.640 W4A8 / INT8 · REF2VA 1,017.780 INT8 / INT8 · REF2VA 1,042.150
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Two dishes: the corrected policy passed the gate Paired frozen two-dish set · nearby starts · candidate v2
Two dishes: the corrected policy passed the gate Frozen baseline: 0 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced. Frozen baseline 0 Frozen baseline: 0 successful trials / 10 Candidate 10 Candidate: 10 successful trials / 10 Gate: 10 0 successful trials / 10
Frozen baseline 0
Candidate 10
successful trials / 10 · gate: 10
One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.
View data & source Two dishes: the corrected policy passed the gate · successful trials / 10 Configuration Value Frozen baseline 0 Candidate 10
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Eight clears still failed the strict gate Mac v3 protocol · paired frozen full-kitchen set
Eight clears still failed the strict gate Frozen baseline: 0 successful trials / 10; Candidate: 8 successful trials / 10. One level and nearby starts, not broad game competence. V3 and v4 are separate protocols. Frozen baseline 0 Frozen baseline: 0 successful trials / 10 Candidate 8 Candidate: 8 successful trials / 10 Gate: 10 0 successful trials / 10
Frozen baseline 0
Candidate 8
successful trials / 10 · gate: 10
One level and nearby starts, not broad game competence. V3 and v4 are separate protocols.
View data & source Eight clears still failed the strict gate · successful trials / 10 Configuration Value Frozen baseline 0 Candidate 8
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The v4 frozen set remained below the gate Mac v4 protocol · different frozen set from v3
The v4 frozen set remained below the gate Frozen baseline: 0 successful trials / 10; Candidate: 7 successful trials / 10. One level and nearby starts, not broad game competence. V3 and v4 are separate protocols. Frozen baseline 0 Frozen baseline: 0 successful trials / 10 Candidate 7 Candidate: 7 successful trials / 10 Gate: 10 0 successful trials / 10
Frozen baseline 0
Candidate 7
successful trials / 10 · gate: 10
One level and nearby starts, not broad game competence. V3 and v4 are separate protocols.
View data & source The v4 frozen set remained below the gate · successful trials / 10 Configuration Value Frozen baseline 0 Candidate 7
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Corrective data vs a same-host retraining control Ubuntu CPU · initial paired full-kitchen evaluation
Corrective data vs a same-host retraining control Retraining control: 7 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced. Retraining control 7 Retraining control: 7 successful trials / 10 Candidate 10 Candidate: 10 successful trials / 10 Gate: 10 0 successful trials / 10
Retraining control 7
Candidate 10
successful trials / 10 · gate: 10
One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.
View data & source Corrective data vs a same-host retraining control · successful trials / 10 Configuration Value Retraining control 7 Candidate 10
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The repeat preserved the candidate result Ubuntu CPU · repeated set · newly retrained control
The repeat preserved the candidate result Retraining control: 5 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. The control had different weights and reached 5/10. Repeating a set adds no new holdout trials. Retraining control 5 Retraining control: 5 successful trials / 10 Candidate 10 Candidate: 10 successful trials / 10 Gate: 10 0 successful trials / 10
Retraining control 5
Candidate 10
successful trials / 10 · gate: 10
One level and nearby starts, not broad game competence. The control had different weights and reached 5/10. Repeating a set adds no new holdout trials.
View data & source The repeat preserved the candidate result · successful trials / 10 Configuration Value Retraining control 5 Candidate 10
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The adapter learned the teacher’s status labels 202 test examples · Qwen3.5-4B QLoRA classifier
The adapter learned the teacher’s status labels Base model: 97 teacher-agreement cases / 202; Tuned model: 192 teacher-agreement cases / 202. Teacher agreement is not independent ground truth. Tuned macro-F1: 0.8625514403292182. Base model 97 Base model: 97 teacher-agreement cases / 202 Tuned model 192 Tuned model: 192 teacher-agreement cases / 202 0 teacher-agreement cases / 202
Base model 97
Tuned model 192
teacher-agreement cases / 202
Teacher agreement is not independent ground truth. Tuned macro-F1: 0.8625514403292182.
View data & source The adapter learned the teacher’s status labels · teacher-agreement cases / 202 Configuration Value Base model 97 Tuned model 192
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Balancing the expert split improved decode GPU expert blocks · 204,800 context · Q8 KV · 300 generated tokens · three RTX 3090s
Balancing the expert split improved decode 8 / 8 / 8: 18.9 tokens / second; 10 / 10 / 10: 22.1 tokens / second; 11 / 11 / 11: 26.1 tokens / second; 12 / 12 / 12: 31.1 tokens / second; 12 / 13 / 12: 30.9 tokens / second. Values transcribed from the report table; cross-runtime comparison uses a different quant. 8 / 8 / 8 18.9 8 / 8 / 8: 18.9 tokens / second 10 / 10 / 10 22.1 10 / 10 / 10: 22.1 tokens / second 11 / 11 / 11 26.1 11 / 11 / 11: 26.1 tokens / second 12 / 12 / 12 31.1 12 / 12 / 12: 31.1 tokens / second 12 / 13 / 12 30.9 12 / 13 / 12: 30.9 tokens / second 0 tokens / second
8 / 8 / 8 18.9
10 / 10 / 10 22.1
11 / 11 / 11 26.1
12 / 12 / 12 31.1
12 / 13 / 12 30.9
tokens / second
Values transcribed from the report table; cross-runtime comparison uses a different quant.
View data & source Balancing the expert split improved decode · tokens / second Configuration Value 8 / 8 / 8 18.9 10 / 10 / 10 22.1 11 / 11 / 11 26.1 12 / 12 / 12 31.1 12 / 13 / 12 30.9
Origin: reported table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result GPU decode shortened the final stage Seed-matched recorded deployment A/B · FP32 reference encoder on CPU
GPU decode shortened the final stage CPU INT8 decode: 3.36 seconds; GPU FP8 decode: 0.24 seconds. Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs. CPU INT8 decode 3.36 CPU INT8 decode: 3.36 seconds GPU FP8 decode 0.24 GPU FP8 decode: 0.24 seconds 0 seconds
CPU INT8 decode 3.36
GPU FP8 decode 0.24
seconds
Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.
View data & source GPU decode shortened the final stage · seconds Configuration Value CPU INT8 decode 3.36 GPU FP8 decode 0.24
Origin: reported table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result End-to-end render wall time Seed-matched recorded deployment A/B · FP32 reference encoder on CPU
End-to-end render wall time Before: 35.10 seconds; After: 24.40 seconds. Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs. Before 35.10 Before: 35.10 seconds After 24.40 After: 24.40 seconds 0 seconds
Before 35.10
After 24.40
seconds
Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.
View data & source End-to-end render wall time · seconds Configuration Value Before 35.10 After 24.40
Origin: reported table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The faster decoder added GPU memory Seed-matched recorded deployment A/B · FP32 reference encoder on CPU
The faster decoder added GPU memory Before: 6.77 GiB; After: 8.01 GiB. Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs. Before 6.77 Before: 6.77 GiB After 8.01 After: 8.01 GiB 0 GiB
Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.
View data & source The faster decoder added GPU memory · GiB Configuration Value Before 6.77 After 8.01
Origin: reported table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Changing the reference encoder recovered similarity AWQ generator · INT8 output decoder · named reference test
Changing the reference encoder recovered similarity AWQ + INT8 reference: 0.8475 WavLM target similarity; AWQ + FP32 reference: 0.9556 WavLM target similarity. A similarity proxy is not a human preference or accent verdict. AWQ + INT8 reference 0.8475 AWQ + INT8 reference: 0.8475 WavLM target similarity AWQ + FP32 reference 0.9556 AWQ + FP32 reference: 0.9556 WavLM target similarity 0 WavLM target similarity
AWQ + INT8 reference 0.8475
AWQ + FP32 reference 0.9556
WavLM target similarity
A similarity proxy is not a human preference or accent verdict.
View data & source Changing the reference encoder recovered similarity · WavLM target similarity Configuration Value AWQ + INT8 reference 0.8475 AWQ + FP32 reference 0.9556
Origin: reported table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Staging moved the peak below 12 GiB Staged-memory run on an RTX 3090
Staging moved the peak below 12 GiB Before staging: 13.01 GiB; After staging: 6.63 GiB. A 12 GiB envelope on a 3090 does not establish physical RTX 3060 acceptance. Before staging 13.01 Before staging: 13.01 GiB After staging 6.63 After staging: 6.63 GiB Gate: 12 0 GiB
Before staging 13.01
After staging 6.63
GiB · gate: 12
A 12 GiB envelope on a 3090 does not establish physical RTX 3060 acceptance.
View data & source Staging moved the peak below 12 GiB · GiB Configuration Value Before staging 13.01 After staging 6.63
Origin: reported table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Thinking-on inbox performance across configurations Same 50 inbox tasks · fixed sampling · different hardware, quantizations and runtimes
Thinking-on inbox performance across configurations A Swift-1.5 27B IQ3_S (Win): 94 strict pass (%); B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 94 strict pass (%); C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 90 strict pass (%); D Gemma4-12B Q4_K_M+MTP (RTX 3060): 76 strict pass (%); E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 32 strict pass (%); F MiMo-9B MLX 4bit (Mac): 52 strict pass (%); F MiMo-9B MLX 4bit (Mac): 64 strict pass (%); G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 10 strict pass (%); H Bonsai-2-27B MLX 2bit (Mac): 94 strict pass (%); I Bonsai-2-27B-Abl PQ2_0 (Mac): 90 strict pass (%); J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 96 strict pass (%); JMac Bonsai-2-27B PQ2_0+MTP (Mac): 96 strict pass (%); L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 92 strict pass (%); M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 82 strict pass (%); K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 96 strict pass (%); KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 98 strict pass (%); NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 86 strict pass (%). One observed workload. Each label retains the hardware; forced thinking is a distinct mode. A Swift-1.5 27B IQ3_S (Win) 94 A Swift-1.5 27B IQ3_S (Win): 94 strict pass (%) B Bonsai-2-27B-Abl PQ2_0 (RTX 3060) 94 B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 94 strict pass (%) C Qwen3.8-27B GSQ-RCO IQ3_S (Win) 90 C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 90 strict pass (%) D Gemma4-12B Q4_K_M+MTP (RTX 3060) 76 D Gemma4-12B Q4_K_M+MTP (RTX 3060): 76 strict pass (%) E Xing4.0-29B-A4B IQ3_XXS (RTX 3060) 32 E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 32 strict pass (%) F MiMo-9B MLX 4bit (Mac) 52 F MiMo-9B MLX 4bit (Mac): 52 strict pass (%) F MiMo-9B MLX 4bit (Mac) 64 F MiMo-9B MLX 4bit (Mac): 64 strict pass (%) G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060) 10 G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 10 strict pass (%) H Bonsai-2-27B MLX 2bit (Mac) 94 H Bonsai-2-27B MLX 2bit (Mac): 94 strict pass (%) I Bonsai-2-27B-Abl PQ2_0 (Mac) 90 I Bonsai-2-27B-Abl PQ2_0 (Mac): 90 strict pass (%) J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060) 96 J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 96 strict pass (%) JMac Bonsai-2-27B PQ2_0+MTP (Mac) 96 JMac Bonsai-2-27B PQ2_0+MTP (Mac): 96 strict pass (%) L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) 92 L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 92 strict pass (%) M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) 82 M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 82 strict pass (%) K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft) 96 K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 96 strict pass (%) KMac Bonsai-2 PQ2_0+DFlash2 (Mac) 98 KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 98 strict pass (%) NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) 86 NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 86 strict pass (%) 0 strict pass (%)
A Swift-1.5 27B IQ3_S (Win) 94
B Bonsai-2-27B-Abl PQ2_0 (RTX 3060) 94
C Qwen3.8-27B GSQ-RCO IQ3_S (Win) 90
D Gemma4-12B Q4_K_M+MTP (RTX 3060) 76
E Xing4.0-29B-A4B IQ3_XXS (RTX 3060) 32
F MiMo-9B MLX 4bit (Mac) 52
F MiMo-9B MLX 4bit (Mac) 64
G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060) 10
H Bonsai-2-27B MLX 2bit (Mac) 94
I Bonsai-2-27B-Abl PQ2_0 (Mac) 90
J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060) 96
JMac Bonsai-2-27B PQ2_0+MTP (Mac) 96
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) 92
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) 82
K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft) 96
KMac Bonsai-2 PQ2_0+DFlash2 (Mac) 98
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) 86
strict pass (%)
One observed workload. Each label retains the hardware; forced thinking is a distinct mode.
View data & source Thinking-on inbox performance across configurations · strict pass (%) Configuration Value A Swift-1.5 27B IQ3_S (Win) 94 B Bonsai-2-27B-Abl PQ2_0 (RTX 3060) 94 C Qwen3.8-27B GSQ-RCO IQ3_S (Win) 90 D Gemma4-12B Q4_K_M+MTP (RTX 3060) 76 E Xing4.0-29B-A4B IQ3_XXS (RTX 3060) 32 F MiMo-9B MLX 4bit (Mac) 52 F MiMo-9B MLX 4bit (Mac) 64 G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060) 10 H Bonsai-2-27B MLX 2bit (Mac) 94 I Bonsai-2-27B-Abl PQ2_0 (Mac) 90 J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060) 96 JMac Bonsai-2-27B PQ2_0+MTP (Mac) 96 L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) 92 M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) 82 K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft) 96 KMac Bonsai-2 PQ2_0+DFlash2 (Mac) 98 NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) 86
Origin: reported aggregate table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The cost of the same inbox workload Thinking-on and forced-thinking rows · same 50 tasks
The cost of the same inbox workload A Swift-1.5 27B IQ3_S (Win): 51.2 median seconds / task; B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 26.6 median seconds / task; C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 59.9 median seconds / task; D Gemma4-12B Q4_K_M+MTP (RTX 3060): 8.0 median seconds / task; E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 69.0 median seconds / task; F MiMo-9B MLX 4bit (Mac): 6.7 median seconds / task; F MiMo-9B MLX 4bit (Mac): 17.4 median seconds / task; G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 94.4 median seconds / task; H Bonsai-2-27B MLX 2bit (Mac): 87.0 median seconds / task; I Bonsai-2-27B-Abl PQ2_0 (Mac): 81.7 median seconds / task; J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 15.8 median seconds / task; JMac Bonsai-2-27B PQ2_0+MTP (Mac): 67.1 median seconds / task; L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 81.9 median seconds / task; M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 65.0 median seconds / task; K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 52.0 median seconds / task; KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 71.6 median seconds / task; NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 135.5 median seconds / task. Median client wall time. A higher pass rate does not automatically give a faster workflow. A Swift-1.5 27B IQ3_S (Win) 51.2 A Swift-1.5 27B IQ3_S (Win): 51.2 median seconds / task B Bonsai-2-27B-Abl PQ2_0 (RTX 3060) 26.6 B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 26.6 median seconds / task C Qwen3.8-27B GSQ-RCO IQ3_S (Win) 59.9 C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 59.9 median seconds / task D Gemma4-12B Q4_K_M+MTP (RTX 3060) 8.0 D Gemma4-12B Q4_K_M+MTP (RTX 3060): 8.0 median seconds / task E Xing4.0-29B-A4B IQ3_XXS (RTX 3060) 69.0 E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 69.0 median seconds / task F MiMo-9B MLX 4bit (Mac) 6.7 F MiMo-9B MLX 4bit (Mac): 6.7 median seconds / task F MiMo-9B MLX 4bit (Mac) 17.4 F MiMo-9B MLX 4bit (Mac): 17.4 median seconds / task G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060) 94.4 G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 94.4 median seconds / task H Bonsai-2-27B MLX 2bit (Mac) 87.0 H Bonsai-2-27B MLX 2bit (Mac): 87.0 median seconds / task I Bonsai-2-27B-Abl PQ2_0 (Mac) 81.7 I Bonsai-2-27B-Abl PQ2_0 (Mac): 81.7 median seconds / task J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060) 15.8 J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 15.8 median seconds / task JMac Bonsai-2-27B PQ2_0+MTP (Mac) 67.1 JMac Bonsai-2-27B PQ2_0+MTP (Mac): 67.1 median seconds / task L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) 81.9 L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 81.9 median seconds / task M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) 65.0 M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 65.0 median seconds / task K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft) 52.0 K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 52.0 median seconds / task KMac Bonsai-2 PQ2_0+DFlash2 (Mac) 71.6 KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 71.6 median seconds / task NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) 135.5 NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 135.5 median seconds / task 0 median seconds / task
A Swift-1.5 27B IQ3_S (Win) 51.2
B Bonsai-2-27B-Abl PQ2_0 (RTX 3060) 26.6
C Qwen3.8-27B GSQ-RCO IQ3_S (Win) 59.9
D Gemma4-12B Q4_K_M+MTP (RTX 3060) 8.0
E Xing4.0-29B-A4B IQ3_XXS (RTX 3060) 69.0
F MiMo-9B MLX 4bit (Mac) 6.7
F MiMo-9B MLX 4bit (Mac) 17.4
G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060) 94.4
H Bonsai-2-27B MLX 2bit (Mac) 87.0
I Bonsai-2-27B-Abl PQ2_0 (Mac) 81.7
J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060) 15.8
JMac Bonsai-2-27B PQ2_0+MTP (Mac) 67.1
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) 81.9
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) 65.0
K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft) 52.0
KMac Bonsai-2 PQ2_0+DFlash2 (Mac) 71.6
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) 135.5
median seconds / task
Median client wall time. A higher pass rate does not automatically give a faster workflow.
View data & source The cost of the same inbox workload · median seconds / task Configuration Value A Swift-1.5 27B IQ3_S (Win) 51.2 B Bonsai-2-27B-Abl PQ2_0 (RTX 3060) 26.6 C Qwen3.8-27B GSQ-RCO IQ3_S (Win) 59.9 D Gemma4-12B Q4_K_M+MTP (RTX 3060) 8.0 E Xing4.0-29B-A4B IQ3_XXS (RTX 3060) 69.0 F MiMo-9B MLX 4bit (Mac) 6.7 F MiMo-9B MLX 4bit (Mac) 17.4 G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060) 94.4 H Bonsai-2-27B MLX 2bit (Mac) 87.0 I Bonsai-2-27B-Abl PQ2_0 (Mac) 81.7 J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060) 15.8 JMac Bonsai-2-27B PQ2_0+MTP (Mac) 67.1 L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) 81.9 M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) 65.0 K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft) 52.0 KMac Bonsai-2 PQ2_0+DFlash2 (Mac) 71.6 NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) 135.5
Origin: reported aggregate table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The quantized classifier retained teacher agreement AutoRound Q4_K_M · 202 frozen teacher-labeled examples · 193/202 total agreement
The quantized classifier retained teacher agreement running: 98.77 recall (%); finished: 80.00 recall (%); needs_input: 85.00 recall (%). Running dominates the set (162 examples); the other two classes have 20 each. No independent ground-truth claim. running 98.77 running: 98.77 recall (%) finished 80.00 finished: 80.00 recall (%) needs_input 85.00 needs_input: 85.00 recall (%) 0 recall (%)
running 98.77
finished 80.00
needs_input 85.00
recall (%)
Running dominates the set (162 examples); the other two classes have 20 each. No independent ground-truth claim.
View data & source The quantized classifier retained teacher agreement · recall (%) Configuration Value running 98.77 finished 80.00 needs_input 85.00
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Codec checkpoint storage, not runtime memory Approximate checkpoint sizes from the conversion notebook
Codec checkpoint storage, not runtime memory FP32: 6.60 decimal GB; BF16: 3.40 decimal GB; INT8 reference: 1.66 decimal GB. The reference INT8 loader materialized BF16 weights at runtime. File size is not a VRAM result. FP32 6.60 FP32: 6.60 decimal GB BF16 3.40 BF16: 3.40 decimal GB INT8 reference 1.66 INT8 reference: 1.66 decimal GB 0 decimal GB
FP32 6.60
BF16 3.40
INT8 reference 1.66
decimal GB
The reference INT8 loader materialized BF16 weights at runtime. File size is not a VRAM result.
View data & source Codec checkpoint storage, not runtime memory · decimal GB Configuration Value FP32 6.60 BF16 3.40 INT8 reference 1.66
Origin: reported conversion sizes. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result First request and warm request are different states HQ4 + INT8 codec lazy service · recorded request timings
First request and warm request are different states Cold first render: 14.8 seconds; Warm render: 5.7 seconds. Cold start includes service wake-up. This is not two equivalent decoding workloads. Cold first render 14.8 Cold first render: 14.8 seconds Warm render 5.7 Warm render: 5.7 seconds 0 seconds
Cold first render 14.8
Warm render 5.7
seconds
Cold start includes service wake-up. This is not two equivalent decoding workloads.
View data & source First request and warm request are different states · seconds Configuration Value Cold first render 14.8 Warm render 5.7
Origin: reported live-service timings. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Sharp reduced completion tokens in all three deployments Medium + xhigh combined · stock vs Sharp v22.4.0
Sharp reduced completion tokens in all three deployments Hauhau Q6: 20.6 completion-token reduction (%); Unsloth Q6: 20.1 completion-token reduction (%); OrcaRouter FP8: 16.0 completion-token reduction (%). The longer prompt reduces the net saving to 9.0%, 6.5% and 3.2%. Small accuracy differences were not significant. Hauhau Q6 20.6 Hauhau Q6: 20.6 completion-token reduction (%) Unsloth Q6 20.1 Unsloth Q6: 20.1 completion-token reduction (%) OrcaRouter FP8 16.0 OrcaRouter FP8: 16.0 completion-token reduction (%) 0 completion-token reduction (%)
Hauhau Q6 20.6
Unsloth Q6 20.1
OrcaRouter FP8 16.0
completion-token reduction (%)
The longer prompt reduces the net saving to 9.0%, 6.5% and 3.2%. Small accuracy differences were not significant.
View data & source Sharp reduced completion tokens in all three deployments · completion-token reduction (%) Configuration Value Hauhau Q6 20.6 Unsloth Q6 20.1 OrcaRouter FP8 16.0
Origin: reported paired aggregate table. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result Fewer observed transcription errors in both languages iPhone 15 Pro · 24 English + 24 Spanish clips · 325 + 242 reference words
Fewer observed transcription errors in both languages English · ONNX · 2 CPU threads: 3.69 word error rate (%); English · Aidana · MLX GPU: 2.15 word error rate (%); Spanish · ONNX · 2 CPU threads: 9.50 word error rate (%); Spanish · Aidana · MLX GPU: 5.37 word error rate (%). Short public read speech, one completed pass per stack. English’s paired interval crosses zero; this does not establish all-accent or noisy-microphone accuracy. English · ONNX · 2 CPU threads 3.69 English · ONNX · 2 CPU threads: 3.69 word error rate (%) English · Aidana · MLX GPU 2.15 English · Aidana · MLX GPU: 2.15 word error rate (%) Spanish · ONNX · 2 CPU threads 9.50 Spanish · ONNX · 2 CPU threads: 9.50 word error rate (%) Spanish · Aidana · MLX GPU 5.37 Spanish · Aidana · MLX GPU: 5.37 word error rate (%) 0 word error rate (%)
English · ONNX · 2 CPU threads 3.69
English · Aidana · MLX GPU 2.15
Spanish · ONNX · 2 CPU threads 9.50
Spanish · Aidana · MLX GPU 5.37
word error rate (%)
Short public read speech, one completed pass per stack. English’s paired interval crosses zero; this does not establish all-accent or noisy-microphone accuracy.
View data & source Fewer observed transcription errors in both languages · word error rate (%) Configuration Value English · ONNX · 2 CPU threads 3.69 English · Aidana · MLX GPU 2.15 Spanish · ONNX · 2 CPU threads 9.50 Spanish · Aidana · MLX GPU 5.37
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The Spanish improvement is clearer on this sample Aidana − ONNX · 5,000 paired clip bootstrap samples · 95% intervals
The Spanish improvement is clearer on this sample English: -1.54 WER difference (percentage points), 95 percent interval -3.49 to 0.29; Spanish: -4.13 WER difference (percentage points), 95 percent interval -8.55 to -0.63. Lower favors Aidana. English crosses zero; the Spanish interval is below zero. This small sample does not establish performance across all speech. English -1.54 95% interval: -3.49 to 0.29 Spanish -4.13 95% interval: -8.55 to -0.63 Zero difference
English -1.54
95% interval: -3.49 to 0.29 · zero marked
Spanish -4.13
95% interval: -8.55 to -0.63 · zero marked
Lower favors Aidana. English crosses zero; the Spanish interval is below zero. This small sample does not establish performance across all speech.
View data & source The Spanish improvement is clearer on this sample · WER difference (percentage points) Configuration Value 95% low 95% high English -1.54 -3.49 0.29 Spanish -4.13 -8.55 -0.63
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON) Recorded result The same 47 clips took less time to transcribe 47 matched clips · 279.724 seconds of audio · iPhone 15 Pro
The same 47 clips took less time to transcribe ONNX · 2 CPU threads: 41.575 decode seconds; Aidana · MLX GPU: 12.400 decode seconds. One background-suspended clip is excluded from both engines. The first-use cost remains: Aidana’s first clip took 2.865 s versus 0.573 s for ONNX. ONNX · 2 CPU threads 41.575 ONNX · 2 CPU threads: 41.575 decode seconds Aidana · MLX GPU 12.400 Aidana · MLX GPU: 12.400 decode seconds 0 decode seconds
ONNX · 2 CPU threads 41.575
Aidana · MLX GPU 12.400
decode seconds
One background-suspended clip is excluded from both engines. The first-use cost remains: Aidana’s first clip took 2.865 s versus 0.573 s for ONNX.
View data & source The same 47 clips took less time to transcribe · decode seconds Configuration Value ONNX · 2 CPU threads 41.575 Aidana · MLX GPU 12.400
Origin: structured measurements. Revision 6550ead3945b.
Download measurement metadata (JSON)