Figures from the workbench

The results.

The measurements, with the conditions still attached.

Every figure comes from a recorded measurement or an explicit report table. Compare within the stated workload, inspect the limits, and take the data with you.

How I report results

Showing 34 figures

Recorded result

What recovery cost in three live scenarios

Single RTX 3090 · local Qwen worker + GLM reviewer · each scenario finished 10/10 checks

What recovery cost in three live scenariosRepeated failures · 2 reviews: 150 seconds; Acceptance failure · 1 review: 142 seconds; Successful task · no review: 9 seconds. Deliberately triggered integration scenarios; duration does not measure general coding accuracy.Repeated failures · 2 reviews150Repeated failures · 2 reviews: 150 secondsAcceptance failure · 1 review142Acceptance failure · 1 review: 142 secondsSuccessful task · no review9Successful task · no review: 9 seconds0seconds
  1. Repeated failures · 2 reviews150
  2. Acceptance failure · 1 review142
  3. Successful task · no review9

seconds

Deliberately triggered integration scenarios; duration does not measure general coding accuracy.

View data & source
What recovery cost in three live scenarios · seconds
ConfigurationValue
Repeated failures · 2 reviews150
Acceptance failure · 1 review142
Successful task · no review9

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

A compact handoff can still add information

Serialized context at the recovery handoff

A compact handoff can still add informationRepeated failures · before: 14,325 characters; Repeated failures · after: 5,723 characters; Acceptance failure · before: 5,290 characters; Acceptance failure · after: 6,566 characters. The short acceptance request grew after adding diagnostic information; this is not a universal compression ratio.Repeated failures · before14,325Repeated failures · before: 14,325 charactersRepeated failures · after5,723Repeated failures · after: 5,723 charactersAcceptance failure · before5,290Acceptance failure · before: 5,290 charactersAcceptance failure · after6,566Acceptance failure · after: 6,566 characters0characters
  1. Repeated failures · before14,325
  2. Repeated failures · after5,723
  3. Acceptance failure · before5,290
  4. Acceptance failure · after6,566

characters

The short acceptance request grew after adding diagnostic information; this is not a universal compression ratio.

View data & source
A compact handoff can still add information · characters
ConfigurationValue
Repeated failures · before14,325
Repeated failures · after5,723
Acceptance failure · before5,290
Acceptance failure · after6,566

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Every candidate crossed the retention floor

Public regression set · first evaluation block · no confirmation/transfer reads

Every candidate crossed the retention floorFrozen parent: 446 checks passed / 446; LR 1e-5 · after 8 steps: 429 checks passed / 446; LR 2e-5 · after 8 steps: 346 checks passed / 446; LR 3e-5 · after 8 steps: 386 checks passed / 446; Restored parent: 446 checks passed / 446. All three runs rolled back. A restored parent is not an improved trained checkpoint.Frozen parent446Frozen parent: 446 checks passed / 446LR 1e-5 · after 8 steps429LR 1e-5 · after 8 steps: 429 checks passed / 446LR 2e-5 · after 8 steps346LR 2e-5 · after 8 steps: 346 checks passed / 446LR 3e-5 · after 8 steps386LR 3e-5 · after 8 steps: 386 checks passed / 446Restored parent446Restored parent: 446 checks passed / 446Gate: 4400checks passed / 446
  1. Frozen parent446
  2. LR 1e-5 · after 8 steps429
  3. LR 2e-5 · after 8 steps346
  4. LR 3e-5 · after 8 steps386
  5. Restored parent446

checks passed / 446 · gate: 440

All three runs rolled back. A restored parent is not an improved trained checkpoint.

View data & source
Every candidate crossed the retention floor · checks passed / 446
ConfigurationValue
Frozen parent446
LR 1e-5 · after 8 steps429
LR 2e-5 · after 8 steps346
LR 3e-5 · after 8 steps386
Restored parent446

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The smaller encoder moved more documents

275 documents · batch size 32 · AWQ W4A16

The smaller encoder moved more documents1B · native 2048d: 1,003.8 documents / second; 8B · native 4096d: 181.0 documents / second. Amortized throughput on this corpus, not single-query latency.1B · native 2048d1,003.81B · native 2048d: 1,003.8 documents / second8B · native 4096d181.08B · native 4096d: 181.0 documents / second0documents / second
  1. 1B · native 2048d1,003.8
  2. 8B · native 4096d181.0

documents / second

Amortized throughput on this corpus, not single-query latency.

View data & source
The smaller encoder moved more documents · documents / second
ConfigurationValue
1B · native 2048d1,003.8
8B · native 4096d181.0

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Memory above the idle GPU baseline

Peak GPU allocation delta · 4 MiB baseline

Memory above the idle GPU baseline1B: 2,380 MiB; 8B: 7,530 MiB. Native dimensions differ. Index storage is a separate cost.1B2,3801B: 2,380 MiB8B7,5308B: 7,530 MiB0MiB
  1. 1B2,380
  2. 8B7,530

MiB

Native dimensions differ. Index storage is a separate cost.

View data & source
Memory above the idle GPU baseline · MiB
ConfigurationValue
1B2,380
8B7,530

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Retrieval scores were close on the bilingual set

110 paired queries · 275 documents · English and Spanish

Retrieval scores were close on the bilingual set1B · native 2048d: 0.955 MRR@10; 8B · native 4096d: 0.962 MRR@10; 8B · sliced 2048d: 0.971 MRR@10. Neither paired bootstrap interval establishes a quality win; see the difference plot.1B · native 2048d0.9551B · native 2048d: 0.955 MRR@108B · native 4096d0.9628B · native 4096d: 0.962 MRR@108B · sliced 2048d0.9718B · sliced 2048d: 0.971 MRR@100MRR@10
  1. 1B · native 2048d0.955
  2. 8B · native 4096d0.962
  3. 8B · sliced 2048d0.971

MRR@10

Neither paired bootstrap interval establishes a quality win; see the difference plot.

View data & source
Retrieval scores were close on the bilingual set · MRR@10
ConfigurationValue
1B · native 2048d0.955
8B · native 4096d0.962
8B · sliced 2048d0.971

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Both quality intervals cross zero

8B − 1B · paired 95% bootstrap intervals · 110 queries

Both quality intervals cross zero8B · sliced_2048d vs 1B: 0.0167 MRR@10 difference, 95 percent interval -0.0061 to 0.0432; 8B · native_4096d vs 1B: 0.0076 MRR@10 difference, 95 percent interval -0.0189 to 0.0371. An interval crossing zero does not establish equivalence or a positive quality gain.8B · sliced_2048d vs 1B0.016795% interval: -0.0061 to 0.04328B · native_4096d vs 1B0.007695% interval: -0.0189 to 0.0371Zero difference
  1. 8B · sliced_2048d vs 1B0.0167

    95% interval: -0.0061 to 0.0432 · zero marked

  2. 8B · native_4096d vs 1B0.0076

    95% interval: -0.0189 to 0.0371 · zero marked

An interval crossing zero does not establish equivalence or a positive quality gain.

View data & source
Both quality intervals cross zero · MRR@10 difference
ConfigurationValue95% low95% high
8B · sliced_2048d vs 1B0.0167-0.00610.0432
8B · native_4096d vs 1B0.0076-0.01890.0371

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Three draft tokens hit the sweet spot

Short-prose mean · EXL3 · Q8 cache · three RTX 3090s

Three draft tokens hit the sweet spotDrafting off: 64.7 tokens / second; 2 draft tokens: 84.4 tokens / second; 3 draft tokens: 88.2 tokens / second; 4 draft tokens: 80.7 tokens / second; 6 draft tokens: 62.4 tokens / second. The separate longer prompt and concurrent streams are different workloads.Drafting off64.7Drafting off: 64.7 tokens / second2 draft tokens84.42 draft tokens: 84.4 tokens / second3 draft tokens88.23 draft tokens: 88.2 tokens / second4 draft tokens80.74 draft tokens: 80.7 tokens / second6 draft tokens62.46 draft tokens: 62.4 tokens / second0tokens / second
  1. Drafting off64.7
  2. 2 draft tokens84.4
  3. 3 draft tokens88.2
  4. 4 draft tokens80.7
  5. 6 draft tokens62.4

tokens / second

The separate longer prompt and concurrent streams are different workloads.

View data & source
Three draft tokens hit the sweet spot · tokens / second
ConfigurationValue
Drafting off64.7
2 draft tokens84.4
3 draft tokens88.2
4 draft tokens80.7
6 draft tokens62.4

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The best checkpoint was not the last one

Five fixed Spanish prompts · 49 reference words · STT screen

The best checkpoint was not the last oneStep 100: 10.2 word error rate (%); Step 200: 14.3 word error rate (%); Step 300: 26.5 word error rate (%); Step 400: 6.1 word error rate (%); Step 500: 42.9 word error rate (%); Step 600: 26.5 word error rate (%); Step 700: 18.4 word error rate (%); Step 800: 24.5 word error rate (%); Step 900: 12.2 word error rate (%); Step 1000: 16.3 word error rate (%); Step 1100: 14.3 word error rate (%); Step 1200: 16.3 word error rate (%); Step 1300: 28.6 word error rate (%); Step 1400: 28.6 word error rate (%); Step 1500: 42.9 word error rate (%). Transcript error does not measure accent, timbre or listener preference.0%12%24%36%48%Step 100: 10.2%Step 200: 14.3%Step 300: 26.5%Step 400: 6.1%Best: 6.1%Step 500: 42.9%Step 600: 26.5%Step 700: 18.4%Step 800: 24.5%Step 900: 12.2%Step 1000: 16.3%Step 1100: 14.3%Step 1200: 16.3%Step 1300: 28.6%Step 1400: 28.6%Step 1500: 42.9%10050010001500Training step · word error rate (%)
0%24%48%Step 100: 10.2%Step 200: 14.3%Step 300: 26.5%Step 400: 6.1%Step 500: 42.9%Step 600: 26.5%Step 700: 18.4%Step 800: 24.5%Step 900: 12.2%Step 1000: 16.3%Step 1100: 14.3%Step 1200: 16.3%Step 1300: 28.6%Step 1400: 28.6%Step 1500: 42.9%10050010001500Training step

Transcript error does not measure accent, timbre or listener preference.

View data & source
The best checkpoint was not the last one · word error rate (%)
ConfigurationValue
Step 10010.2
Step 20014.3
Step 30026.5
Step 4006.1
Step 50042.9
Step 60026.5
Step 70018.4
Step 80024.5
Step 90012.2
Step 100016.3
Step 110014.3
Step 120016.3
Step 130028.6
Step 140028.6
Step 150042.9

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Matched scenes, different diffusion precision

1024×1024 · 40 steps · matched seeds 2401–2408 · warm RTX 3090 runs

Matched scenes, different diffusion precisionFP8 · mean of 8 scenes: 93.352 seconds; INT8 · mean of 8 scenes: 46.219 seconds. INT8 was opt-in. Fine detail changed, so timing improvement is not visual-quality parity.FP8 · mean of 8 scenes93.352FP8 · mean of 8 scenes: 93.352 secondsINT8 · mean of 8 scenes46.219INT8 · mean of 8 scenes: 46.219 seconds0seconds
  1. FP8 · mean of 8 scenes93.352
  2. INT8 · mean of 8 scenes46.219

seconds

INT8 was opt-in. Fine detail changed, so timing improvement is not visual-quality parity.

View data & source
Matched scenes, different diffusion precision · seconds
ConfigurationValue
FP8 · mean of 8 scenes93.352
INT8 · mean of 8 scenes46.219

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Every matched scene in the timing suite

Same eight scenes · FP8 and INT8 · 40 steps

Every matched scene in the timing suiteSeed 2401 · FP8: 93.215 seconds; Seed 2401 · INT8: 46.079 seconds; Seed 2402 · FP8: 93.244 seconds; Seed 2402 · INT8: 46.154 seconds; Seed 2403 · FP8: 93.446 seconds; Seed 2403 · INT8: 46.262 seconds; Seed 2404 · FP8: 93.395 seconds; Seed 2404 · INT8: 46.265 seconds; Seed 2405 · FP8: 93.318 seconds; Seed 2405 · INT8: 46.217 seconds; Seed 2406 · FP8: 93.357 seconds; Seed 2406 · INT8: 46.271 seconds; Seed 2407 · FP8: 93.449 seconds; Seed 2407 · INT8: 46.274 seconds; Seed 2408 · FP8: 93.389 seconds; Seed 2408 · INT8: 46.227 seconds. Each pair shares a seed and workload. Error bars are unavailable; these are recorded runs.Seed 2401 · FP893.215Seed 2401 · FP8: 93.215 secondsSeed 2401 · INT846.079Seed 2401 · INT8: 46.079 secondsSeed 2402 · FP893.244Seed 2402 · FP8: 93.244 secondsSeed 2402 · INT846.154Seed 2402 · INT8: 46.154 secondsSeed 2403 · FP893.446Seed 2403 · FP8: 93.446 secondsSeed 2403 · INT846.262Seed 2403 · INT8: 46.262 secondsSeed 2404 · FP893.395Seed 2404 · FP8: 93.395 secondsSeed 2404 · INT846.265Seed 2404 · INT8: 46.265 secondsSeed 2405 · FP893.318Seed 2405 · FP8: 93.318 secondsSeed 2405 · INT846.217Seed 2405 · INT8: 46.217 secondsSeed 2406 · FP893.357Seed 2406 · FP8: 93.357 secondsSeed 2406 · INT846.271Seed 2406 · INT8: 46.271 secondsSeed 2407 · FP893.449Seed 2407 · FP8: 93.449 secondsSeed 2407 · INT846.274Seed 2407 · INT8: 46.274 secondsSeed 2408 · FP893.389Seed 2408 · FP8: 93.389 secondsSeed 2408 · INT846.227Seed 2408 · INT8: 46.227 seconds0seconds
  1. Seed 2401 · FP893.215
  2. Seed 2401 · INT846.079
  3. Seed 2402 · FP893.244
  4. Seed 2402 · INT846.154
  5. Seed 2403 · FP893.446
  6. Seed 2403 · INT846.262
  7. Seed 2404 · FP893.395
  8. Seed 2404 · INT846.265
  9. Seed 2405 · FP893.318
  10. Seed 2405 · INT846.217
  11. Seed 2406 · FP893.357
  12. Seed 2406 · INT846.271
  13. Seed 2407 · FP893.449
  14. Seed 2407 · INT846.274
  15. Seed 2408 · FP893.389
  16. Seed 2408 · INT846.227

seconds

Each pair shares a seed and workload. Error bars are unavailable; these are recorded runs.

View data & source
Every matched scene in the timing suite · seconds
ConfigurationValue
Seed 2401 · FP893.215
Seed 2401 · INT846.079
Seed 2402 · FP893.244
Seed 2402 · INT846.154
Seed 2403 · FP893.446
Seed 2403 · INT846.262
Seed 2404 · FP893.395
Seed 2404 · INT846.265
Seed 2405 · FP893.318
Seed 2405 · INT846.217
Seed 2406 · FP893.357
Seed 2406 · INT846.271
Seed 2407 · FP893.449
Seed 2407 · INT846.274
Seed 2408 · FP893.389
Seed 2408 · INT846.227

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

A five-second video at 960×544

Three RTX 3090s · 124 frames · 24 fps · 20 steps · full-decode checks passed

A five-second video at 960×544W4A8 / NVFP4 · REF2VA: 319.045 elapsed seconds; W4A8 / INT8 · multishot: 321.739 elapsed seconds; W4A8 / NVFP4 · REF2VA · alt. placement: 331.366 elapsed seconds; INT8 / INT8 · T2VA/FL2VA: 354.580 elapsed seconds; INT4Q / INT8 · multishot: 380.578 elapsed seconds; INT8 / INT8 · REF2VA · HQ: 399.403 elapsed seconds; BF16 / Q4_K_M · REF2VA: 501.095 elapsed seconds. Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.W4A8 / NVFP4 · REF2VA319.045W4A8 / NVFP4 · REF2VA: 319.045 elapsed secondsW4A8 / INT8 · multishot321.739W4A8 / INT8 · multishot: 321.739 elapsed secondsW4A8 / NVFP4 · REF2VA · alt. placement331.366W4A8 / NVFP4 · REF2VA · alt. placement: 331.366 elapsed secondsINT8 / INT8 · T2VA/FL2VA354.580INT8 / INT8 · T2VA/FL2VA: 354.580 elapsed secondsINT4Q / INT8 · multishot380.578INT4Q / INT8 · multishot: 380.578 elapsed secondsINT8 / INT8 · REF2VA · HQ399.403INT8 / INT8 · REF2VA · HQ: 399.403 elapsed secondsBF16 / Q4_K_M · REF2VA501.095BF16 / Q4_K_M · REF2VA: 501.095 elapsed seconds0elapsed seconds
  1. W4A8 / NVFP4 · REF2VA319.045
  2. W4A8 / INT8 · multishot321.739
  3. W4A8 / NVFP4 · REF2VA · alt. placement331.366
  4. INT8 / INT8 · T2VA/FL2VA354.580
  5. INT4Q / INT8 · multishot380.578
  6. INT8 / INT8 · REF2VA · HQ399.403
  7. BF16 / Q4_K_M · REF2VA501.095

elapsed seconds

Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.

View data & source
A five-second video at 960×544 · elapsed seconds
ConfigurationValue
W4A8 / NVFP4 · REF2VA319.045
W4A8 / INT8 · multishot321.739
W4A8 / NVFP4 · REF2VA · alt. placement331.366
INT8 / INT8 · T2VA/FL2VA354.580
INT4Q / INT8 · multishot380.578
INT8 / INT8 · REF2VA · HQ399.403
BF16 / Q4_K_M · REF2VA501.095

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

A five-second video at 1344×768

Three RTX 3090s · 124 frames · 24 fps · 20 steps · full-decode checks passed

A five-second video at 1344×768W4A8 / W4A8 · REF2VA: 970.306 elapsed seconds; INT8 / INT8 · T2VA/FL2VA: 990.640 elapsed seconds; W4A8 / INT8 · REF2VA: 1,017.780 elapsed seconds; INT8 / INT8 · REF2VA: 1,042.150 elapsed seconds. Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.W4A8 / W4A8 · REF2VA970.306W4A8 / W4A8 · REF2VA: 970.306 elapsed secondsINT8 / INT8 · T2VA/FL2VA990.640INT8 / INT8 · T2VA/FL2VA: 990.640 elapsed secondsW4A8 / INT8 · REF2VA1,017.780W4A8 / INT8 · REF2VA: 1,017.780 elapsed secondsINT8 / INT8 · REF2VA1,042.150INT8 / INT8 · REF2VA: 1,042.150 elapsed seconds0elapsed seconds
  1. W4A8 / W4A8 · REF2VA970.306
  2. INT8 / INT8 · T2VA/FL2VA990.640
  3. W4A8 / INT8 · REF2VA1,017.780
  4. INT8 / INT8 · REF2VA1,042.150

elapsed seconds

Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.

View data & source
A five-second video at 1344×768 · elapsed seconds
ConfigurationValue
W4A8 / W4A8 · REF2VA970.306
INT8 / INT8 · T2VA/FL2VA990.640
W4A8 / INT8 · REF2VA1,017.780
INT8 / INT8 · REF2VA1,042.150

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Two dishes: the corrected policy passed the gate

Paired frozen two-dish set · nearby starts · candidate v2

Two dishes: the corrected policy passed the gateFrozen baseline: 0 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.Frozen baseline0Frozen baseline: 0 successful trials / 10Candidate10Candidate: 10 successful trials / 10Gate: 100successful trials / 10
  1. Frozen baseline0
  2. Candidate10

successful trials / 10 · gate: 10

One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.

View data & source
Two dishes: the corrected policy passed the gate · successful trials / 10
ConfigurationValue
Frozen baseline0
Candidate10

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Eight clears still failed the strict gate

Mac v3 protocol · paired frozen full-kitchen set

Eight clears still failed the strict gateFrozen baseline: 0 successful trials / 10; Candidate: 8 successful trials / 10. One level and nearby starts, not broad game competence. V3 and v4 are separate protocols.Frozen baseline0Frozen baseline: 0 successful trials / 10Candidate8Candidate: 8 successful trials / 10Gate: 100successful trials / 10
  1. Frozen baseline0
  2. Candidate8

successful trials / 10 · gate: 10

One level and nearby starts, not broad game competence. V3 and v4 are separate protocols.

View data & source
Eight clears still failed the strict gate · successful trials / 10
ConfigurationValue
Frozen baseline0
Candidate8

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The v4 frozen set remained below the gate

Mac v4 protocol · different frozen set from v3

The v4 frozen set remained below the gateFrozen baseline: 0 successful trials / 10; Candidate: 7 successful trials / 10. One level and nearby starts, not broad game competence. V3 and v4 are separate protocols.Frozen baseline0Frozen baseline: 0 successful trials / 10Candidate7Candidate: 7 successful trials / 10Gate: 100successful trials / 10
  1. Frozen baseline0
  2. Candidate7

successful trials / 10 · gate: 10

One level and nearby starts, not broad game competence. V3 and v4 are separate protocols.

View data & source
The v4 frozen set remained below the gate · successful trials / 10
ConfigurationValue
Frozen baseline0
Candidate7

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Corrective data vs a same-host retraining control

Ubuntu CPU · initial paired full-kitchen evaluation

Corrective data vs a same-host retraining controlRetraining control: 7 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.Retraining control7Retraining control: 7 successful trials / 10Candidate10Candidate: 10 successful trials / 10Gate: 100successful trials / 10
  1. Retraining control7
  2. Candidate10

successful trials / 10 · gate: 10

One level and nearby starts, not broad game competence. All recorded traces replayed exactly; the incumbent was not replaced.

View data & source
Corrective data vs a same-host retraining control · successful trials / 10
ConfigurationValue
Retraining control7
Candidate10

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The repeat preserved the candidate result

Ubuntu CPU · repeated set · newly retrained control

The repeat preserved the candidate resultRetraining control: 5 successful trials / 10; Candidate: 10 successful trials / 10. One level and nearby starts, not broad game competence. The control had different weights and reached 5/10. Repeating a set adds no new holdout trials.Retraining control5Retraining control: 5 successful trials / 10Candidate10Candidate: 10 successful trials / 10Gate: 100successful trials / 10
  1. Retraining control5
  2. Candidate10

successful trials / 10 · gate: 10

One level and nearby starts, not broad game competence. The control had different weights and reached 5/10. Repeating a set adds no new holdout trials.

View data & source
The repeat preserved the candidate result · successful trials / 10
ConfigurationValue
Retraining control5
Candidate10

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The adapter learned the teacher’s status labels

202 test examples · Qwen3.5-4B QLoRA classifier

The adapter learned the teacher’s status labelsBase model: 97 teacher-agreement cases / 202; Tuned model: 192 teacher-agreement cases / 202. Teacher agreement is not independent ground truth. Tuned macro-F1: 0.8625514403292182.Base model97Base model: 97 teacher-agreement cases / 202Tuned model192Tuned model: 192 teacher-agreement cases / 2020teacher-agreement cases / 202
  1. Base model97
  2. Tuned model192

teacher-agreement cases / 202

Teacher agreement is not independent ground truth. Tuned macro-F1: 0.8625514403292182.

View data & source
The adapter learned the teacher’s status labels · teacher-agreement cases / 202
ConfigurationValue
Base model97
Tuned model192

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Balancing the expert split improved decode

GPU expert blocks · 204,800 context · Q8 KV · 300 generated tokens · three RTX 3090s

Balancing the expert split improved decode8 / 8 / 8: 18.9 tokens / second; 10 / 10 / 10: 22.1 tokens / second; 11 / 11 / 11: 26.1 tokens / second; 12 / 12 / 12: 31.1 tokens / second; 12 / 13 / 12: 30.9 tokens / second. Values transcribed from the report table; cross-runtime comparison uses a different quant.8 / 8 / 818.98 / 8 / 8: 18.9 tokens / second10 / 10 / 1022.110 / 10 / 10: 22.1 tokens / second11 / 11 / 1126.111 / 11 / 11: 26.1 tokens / second12 / 12 / 1231.112 / 12 / 12: 31.1 tokens / second12 / 13 / 1230.912 / 13 / 12: 30.9 tokens / second0tokens / second
  1. 8 / 8 / 818.9
  2. 10 / 10 / 1022.1
  3. 11 / 11 / 1126.1
  4. 12 / 12 / 1231.1
  5. 12 / 13 / 1230.9

tokens / second

Values transcribed from the report table; cross-runtime comparison uses a different quant.

View data & source
Balancing the expert split improved decode · tokens / second
ConfigurationValue
8 / 8 / 818.9
10 / 10 / 1022.1
11 / 11 / 1126.1
12 / 12 / 1231.1
12 / 13 / 1230.9

Origin: reported table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

GPU decode shortened the final stage

Seed-matched recorded deployment A/B · FP32 reference encoder on CPU

GPU decode shortened the final stageCPU INT8 decode: 3.36 seconds; GPU FP8 decode: 0.24 seconds. Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.CPU INT8 decode3.36CPU INT8 decode: 3.36 secondsGPU FP8 decode0.24GPU FP8 decode: 0.24 seconds0seconds
  1. CPU INT8 decode3.36
  2. GPU FP8 decode0.24

seconds

Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.

View data & source
GPU decode shortened the final stage · seconds
ConfigurationValue
CPU INT8 decode3.36
GPU FP8 decode0.24

Origin: reported table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

End-to-end render wall time

Seed-matched recorded deployment A/B · FP32 reference encoder on CPU

End-to-end render wall timeBefore: 35.10 seconds; After: 24.40 seconds. Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.Before35.10Before: 35.10 secondsAfter24.40After: 24.40 seconds0seconds
  1. Before35.10
  2. After24.40

seconds

Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.

View data & source
End-to-end render wall time · seconds
ConfigurationValue
Before35.10
After24.40

Origin: reported table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The faster decoder added GPU memory

Seed-matched recorded deployment A/B · FP32 reference encoder on CPU

The faster decoder added GPU memoryBefore: 6.77 GiB; After: 8.01 GiB. Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.Before6.77Before: 6.77 GiBAfter8.01After: 8.01 GiB0GiB
  1. Before6.77
  2. After8.01

GiB

Report-table values. The archived before JSON is empty; these are not presented as recovered raw logs.

View data & source
The faster decoder added GPU memory · GiB
ConfigurationValue
Before6.77
After8.01

Origin: reported table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Changing the reference encoder recovered similarity

AWQ generator · INT8 output decoder · named reference test

Changing the reference encoder recovered similarityAWQ + INT8 reference: 0.8475 WavLM target similarity; AWQ + FP32 reference: 0.9556 WavLM target similarity. A similarity proxy is not a human preference or accent verdict.AWQ + INT8 reference0.8475AWQ + INT8 reference: 0.8475 WavLM target similarityAWQ + FP32 reference0.9556AWQ + FP32 reference: 0.9556 WavLM target similarity0WavLM target similarity
  1. AWQ + INT8 reference0.8475
  2. AWQ + FP32 reference0.9556

WavLM target similarity

A similarity proxy is not a human preference or accent verdict.

View data & source
Changing the reference encoder recovered similarity · WavLM target similarity
ConfigurationValue
AWQ + INT8 reference0.8475
AWQ + FP32 reference0.9556

Origin: reported table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Staging moved the peak below 12 GiB

Staged-memory run on an RTX 3090

Staging moved the peak below 12 GiBBefore staging: 13.01 GiB; After staging: 6.63 GiB. A 12 GiB envelope on a 3090 does not establish physical RTX 3060 acceptance.Before staging13.01Before staging: 13.01 GiBAfter staging6.63After staging: 6.63 GiBGate: 120GiB
  1. Before staging13.01
  2. After staging6.63

GiB · gate: 12

A 12 GiB envelope on a 3090 does not establish physical RTX 3060 acceptance.

View data & source
Staging moved the peak below 12 GiB · GiB
ConfigurationValue
Before staging13.01
After staging6.63

Origin: reported table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Thinking-on inbox performance across configurations

Same 50 inbox tasks · fixed sampling · different hardware, quantizations and runtimes

Thinking-on inbox performance across configurationsA Swift-1.5 27B IQ3_S (Win): 94 strict pass (%); B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 94 strict pass (%); C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 90 strict pass (%); D Gemma4-12B Q4_K_M+MTP (RTX 3060): 76 strict pass (%); E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 32 strict pass (%); F MiMo-9B MLX 4bit (Mac): 52 strict pass (%); F MiMo-9B MLX 4bit (Mac): 64 strict pass (%); G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 10 strict pass (%); H Bonsai-2-27B MLX 2bit (Mac): 94 strict pass (%); I Bonsai-2-27B-Abl PQ2_0 (Mac): 90 strict pass (%); J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 96 strict pass (%); JMac Bonsai-2-27B PQ2_0+MTP (Mac): 96 strict pass (%); L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 92 strict pass (%); M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 82 strict pass (%); K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 96 strict pass (%); KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 98 strict pass (%); NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 86 strict pass (%). One observed workload. Each label retains the hardware; forced thinking is a distinct mode.A Swift-1.5 27B IQ3_S (Win)94A Swift-1.5 27B IQ3_S (Win): 94 strict pass (%)B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)94B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 94 strict pass (%)C Qwen3.8-27B GSQ-RCO IQ3_S (Win)90C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 90 strict pass (%)D Gemma4-12B Q4_K_M+MTP (RTX 3060)76D Gemma4-12B Q4_K_M+MTP (RTX 3060): 76 strict pass (%)E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)32E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 32 strict pass (%)F MiMo-9B MLX 4bit (Mac)52F MiMo-9B MLX 4bit (Mac): 52 strict pass (%)F MiMo-9B MLX 4bit (Mac)64F MiMo-9B MLX 4bit (Mac): 64 strict pass (%)G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)10G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 10 strict pass (%)H Bonsai-2-27B MLX 2bit (Mac)94H Bonsai-2-27B MLX 2bit (Mac): 94 strict pass (%)I Bonsai-2-27B-Abl PQ2_0 (Mac)90I Bonsai-2-27B-Abl PQ2_0 (Mac): 90 strict pass (%)J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)96J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 96 strict pass (%)JMac Bonsai-2-27B PQ2_0+MTP (Mac)96JMac Bonsai-2-27B PQ2_0+MTP (Mac): 96 strict pass (%)L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)92L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 92 strict pass (%)M ZDTaichu5.0-9B MLX 8bit (Mac,text-only)82M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 82 strict pass (%)K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPUdraft)96K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 96 strict pass (%)KMac Bonsai-2 PQ2_0+DFlash2 (Mac)98KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 98 strict pass (%)NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP(Mac)86NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 86 strict pass (%)0strict pass (%)
  1. A Swift-1.5 27B IQ3_S (Win)94
  2. B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)94
  3. C Qwen3.8-27B GSQ-RCO IQ3_S (Win)90
  4. D Gemma4-12B Q4_K_M+MTP (RTX 3060)76
  5. E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)32
  6. F MiMo-9B MLX 4bit (Mac)52
  7. F MiMo-9B MLX 4bit (Mac)64
  8. G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)10
  9. H Bonsai-2-27B MLX 2bit (Mac)94
  10. I Bonsai-2-27B-Abl PQ2_0 (Mac)90
  11. J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)96
  12. JMac Bonsai-2-27B PQ2_0+MTP (Mac)96
  13. L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)92
  14. M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)82
  15. K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)96
  16. KMac Bonsai-2 PQ2_0+DFlash2 (Mac)98
  17. NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)86

strict pass (%)

One observed workload. Each label retains the hardware; forced thinking is a distinct mode.

View data & source
Thinking-on inbox performance across configurations · strict pass (%)
ConfigurationValue
A Swift-1.5 27B IQ3_S (Win)94
B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)94
C Qwen3.8-27B GSQ-RCO IQ3_S (Win)90
D Gemma4-12B Q4_K_M+MTP (RTX 3060)76
E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)32
F MiMo-9B MLX 4bit (Mac)52
F MiMo-9B MLX 4bit (Mac)64
G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)10
H Bonsai-2-27B MLX 2bit (Mac)94
I Bonsai-2-27B-Abl PQ2_0 (Mac)90
J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)96
JMac Bonsai-2-27B PQ2_0+MTP (Mac)96
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)92
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)82
K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)96
KMac Bonsai-2 PQ2_0+DFlash2 (Mac)98
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)86

Origin: reported aggregate table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The cost of the same inbox workload

Thinking-on and forced-thinking rows · same 50 tasks

The cost of the same inbox workloadA Swift-1.5 27B IQ3_S (Win): 51.2 median seconds / task; B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 26.6 median seconds / task; C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 59.9 median seconds / task; D Gemma4-12B Q4_K_M+MTP (RTX 3060): 8.0 median seconds / task; E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 69.0 median seconds / task; F MiMo-9B MLX 4bit (Mac): 6.7 median seconds / task; F MiMo-9B MLX 4bit (Mac): 17.4 median seconds / task; G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 94.4 median seconds / task; H Bonsai-2-27B MLX 2bit (Mac): 87.0 median seconds / task; I Bonsai-2-27B-Abl PQ2_0 (Mac): 81.7 median seconds / task; J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 15.8 median seconds / task; JMac Bonsai-2-27B PQ2_0+MTP (Mac): 67.1 median seconds / task; L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 81.9 median seconds / task; M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 65.0 median seconds / task; K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 52.0 median seconds / task; KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 71.6 median seconds / task; NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 135.5 median seconds / task. Median client wall time. A higher pass rate does not automatically give a faster workflow.A Swift-1.5 27B IQ3_S (Win)51.2A Swift-1.5 27B IQ3_S (Win): 51.2 median seconds / taskB Bonsai-2-27B-Abl PQ2_0 (RTX 3060)26.6B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 26.6 median seconds / taskC Qwen3.8-27B GSQ-RCO IQ3_S (Win)59.9C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 59.9 median seconds / taskD Gemma4-12B Q4_K_M+MTP (RTX 3060)8.0D Gemma4-12B Q4_K_M+MTP (RTX 3060): 8.0 median seconds / taskE Xing4.0-29B-A4B IQ3_XXS (RTX 3060)69.0E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 69.0 median seconds / taskF MiMo-9B MLX 4bit (Mac)6.7F MiMo-9B MLX 4bit (Mac): 6.7 median seconds / taskF MiMo-9B MLX 4bit (Mac)17.4F MiMo-9B MLX 4bit (Mac): 17.4 median seconds / taskG Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)94.4G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 94.4 median seconds / taskH Bonsai-2-27B MLX 2bit (Mac)87.0H Bonsai-2-27B MLX 2bit (Mac): 87.0 median seconds / taskI Bonsai-2-27B-Abl PQ2_0 (Mac)81.7I Bonsai-2-27B-Abl PQ2_0 (Mac): 81.7 median seconds / taskJ54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)15.8J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 15.8 median seconds / taskJMac Bonsai-2-27B PQ2_0+MTP (Mac)67.1JMac Bonsai-2-27B PQ2_0+MTP (Mac): 67.1 median seconds / taskL ThinkingCap-Qwen3.8-27B IQ3_S (Mac)81.9L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 81.9 median seconds / taskM ZDTaichu5.0-9B MLX 8bit (Mac,text-only)65.0M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 65.0 median seconds / taskK54 Bonsai-2 PQ2_0+DFlash2 (3060, CPUdraft)52.0K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 52.0 median seconds / taskKMac Bonsai-2 PQ2_0+DFlash2 (Mac)71.6KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 71.6 median seconds / taskNMac HauhauCS Qwen3.8-27B IQ3_XS+MTP(Mac)135.5NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 135.5 median seconds / task0median seconds / task
  1. A Swift-1.5 27B IQ3_S (Win)51.2
  2. B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)26.6
  3. C Qwen3.8-27B GSQ-RCO IQ3_S (Win)59.9
  4. D Gemma4-12B Q4_K_M+MTP (RTX 3060)8.0
  5. E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)69.0
  6. F MiMo-9B MLX 4bit (Mac)6.7
  7. F MiMo-9B MLX 4bit (Mac)17.4
  8. G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)94.4
  9. H Bonsai-2-27B MLX 2bit (Mac)87.0
  10. I Bonsai-2-27B-Abl PQ2_0 (Mac)81.7
  11. J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)15.8
  12. JMac Bonsai-2-27B PQ2_0+MTP (Mac)67.1
  13. L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)81.9
  14. M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)65.0
  15. K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)52.0
  16. KMac Bonsai-2 PQ2_0+DFlash2 (Mac)71.6
  17. NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)135.5

median seconds / task

Median client wall time. A higher pass rate does not automatically give a faster workflow.

View data & source
The cost of the same inbox workload · median seconds / task
ConfigurationValue
A Swift-1.5 27B IQ3_S (Win)51.2
B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)26.6
C Qwen3.8-27B GSQ-RCO IQ3_S (Win)59.9
D Gemma4-12B Q4_K_M+MTP (RTX 3060)8.0
E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)69.0
F MiMo-9B MLX 4bit (Mac)6.7
F MiMo-9B MLX 4bit (Mac)17.4
G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)94.4
H Bonsai-2-27B MLX 2bit (Mac)87.0
I Bonsai-2-27B-Abl PQ2_0 (Mac)81.7
J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)15.8
JMac Bonsai-2-27B PQ2_0+MTP (Mac)67.1
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)81.9
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)65.0
K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)52.0
KMac Bonsai-2 PQ2_0+DFlash2 (Mac)71.6
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)135.5

Origin: reported aggregate table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The quantized classifier retained teacher agreement

AutoRound Q4_K_M · 202 frozen teacher-labeled examples · 193/202 total agreement

The quantized classifier retained teacher agreementrunning: 98.77 recall (%); finished: 80.00 recall (%); needs_input: 85.00 recall (%). Running dominates the set (162 examples); the other two classes have 20 each. No independent ground-truth claim.running98.77running: 98.77 recall (%)finished80.00finished: 80.00 recall (%)needs_input85.00needs_input: 85.00 recall (%)0recall (%)
  1. running98.77
  2. finished80.00
  3. needs_input85.00

recall (%)

Running dominates the set (162 examples); the other two classes have 20 each. No independent ground-truth claim.

View data & source
The quantized classifier retained teacher agreement · recall (%)
ConfigurationValue
running98.77
finished80.00
needs_input85.00

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Codec checkpoint storage, not runtime memory

Approximate checkpoint sizes from the conversion notebook

Codec checkpoint storage, not runtime memoryFP32: 6.60 decimal GB; BF16: 3.40 decimal GB; INT8 reference: 1.66 decimal GB. The reference INT8 loader materialized BF16 weights at runtime. File size is not a VRAM result.FP326.60FP32: 6.60 decimal GBBF163.40BF16: 3.40 decimal GBINT8 reference1.66INT8 reference: 1.66 decimal GB0decimal GB
  1. FP326.60
  2. BF163.40
  3. INT8 reference1.66

decimal GB

The reference INT8 loader materialized BF16 weights at runtime. File size is not a VRAM result.

View data & source
Codec checkpoint storage, not runtime memory · decimal GB
ConfigurationValue
FP326.60
BF163.40
INT8 reference1.66

Origin: reported conversion sizes. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

First request and warm request are different states

HQ4 + INT8 codec lazy service · recorded request timings

First request and warm request are different statesCold first render: 14.8 seconds; Warm render: 5.7 seconds. Cold start includes service wake-up. This is not two equivalent decoding workloads.Cold first render14.8Cold first render: 14.8 secondsWarm render5.7Warm render: 5.7 seconds0seconds
  1. Cold first render14.8
  2. Warm render5.7

seconds

Cold start includes service wake-up. This is not two equivalent decoding workloads.

View data & source
First request and warm request are different states · seconds
ConfigurationValue
Cold first render14.8
Warm render5.7

Origin: reported live-service timings. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Sharp reduced completion tokens in all three deployments

Medium + xhigh combined · stock vs Sharp v22.4.0

Sharp reduced completion tokens in all three deploymentsHauhau Q6: 20.6 completion-token reduction (%); Unsloth Q6: 20.1 completion-token reduction (%); OrcaRouter FP8: 16.0 completion-token reduction (%). The longer prompt reduces the net saving to 9.0%, 6.5% and 3.2%. Small accuracy differences were not significant.Hauhau Q620.6Hauhau Q6: 20.6 completion-token reduction (%)Unsloth Q620.1Unsloth Q6: 20.1 completion-token reduction (%)OrcaRouter FP816.0OrcaRouter FP8: 16.0 completion-token reduction (%)0completion-token reduction (%)
  1. Hauhau Q620.6
  2. Unsloth Q620.1
  3. OrcaRouter FP816.0

completion-token reduction (%)

The longer prompt reduces the net saving to 9.0%, 6.5% and 3.2%. Small accuracy differences were not significant.

View data & source
Sharp reduced completion tokens in all three deployments · completion-token reduction (%)
ConfigurationValue
Hauhau Q620.6
Unsloth Q620.1
OrcaRouter FP816.0

Origin: reported paired aggregate table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

Fewer observed transcription errors in both languages

iPhone 15 Pro · 24 English + 24 Spanish clips · 325 + 242 reference words

Fewer observed transcription errors in both languagesEnglish · ONNX · 2 CPU threads: 3.69 word error rate (%); English · Aidana · MLX GPU: 2.15 word error rate (%); Spanish · ONNX · 2 CPU threads: 9.50 word error rate (%); Spanish · Aidana · MLX GPU: 5.37 word error rate (%). Short public read speech, one completed pass per stack. English’s paired interval crosses zero; this does not establish all-accent or noisy-microphone accuracy.English · ONNX · 2 CPU threads3.69English · ONNX · 2 CPU threads: 3.69 word error rate (%)English · Aidana · MLX GPU2.15English · Aidana · MLX GPU: 2.15 word error rate (%)Spanish · ONNX · 2 CPU threads9.50Spanish · ONNX · 2 CPU threads: 9.50 word error rate (%)Spanish · Aidana · MLX GPU5.37Spanish · Aidana · MLX GPU: 5.37 word error rate (%)0word error rate (%)
  1. English · ONNX · 2 CPU threads3.69
  2. English · Aidana · MLX GPU2.15
  3. Spanish · ONNX · 2 CPU threads9.50
  4. Spanish · Aidana · MLX GPU5.37

word error rate (%)

Short public read speech, one completed pass per stack. English’s paired interval crosses zero; this does not establish all-accent or noisy-microphone accuracy.

View data & source
Fewer observed transcription errors in both languages · word error rate (%)
ConfigurationValue
English · ONNX · 2 CPU threads3.69
English · Aidana · MLX GPU2.15
Spanish · ONNX · 2 CPU threads9.50
Spanish · Aidana · MLX GPU5.37

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The Spanish improvement is clearer on this sample

Aidana − ONNX · 5,000 paired clip bootstrap samples · 95% intervals

The Spanish improvement is clearer on this sampleEnglish: -1.54 WER difference (percentage points), 95 percent interval -3.49 to 0.29; Spanish: -4.13 WER difference (percentage points), 95 percent interval -8.55 to -0.63. Lower favors Aidana. English crosses zero; the Spanish interval is below zero. This small sample does not establish performance across all speech.English-1.5495% interval: -3.49 to 0.29Spanish-4.1395% interval: -8.55 to -0.63Zero difference
  1. English-1.54

    95% interval: -3.49 to 0.29 · zero marked

  2. Spanish-4.13

    95% interval: -8.55 to -0.63 · zero marked

Lower favors Aidana. English crosses zero; the Spanish interval is below zero. This small sample does not establish performance across all speech.

View data & source
The Spanish improvement is clearer on this sample · WER difference (percentage points)
ConfigurationValue95% low95% high
English-1.54-3.490.29
Spanish-4.13-8.55-0.63

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The same 47 clips took less time to transcribe

47 matched clips · 279.724 seconds of audio · iPhone 15 Pro

The same 47 clips took less time to transcribeONNX · 2 CPU threads: 41.575 decode seconds; Aidana · MLX GPU: 12.400 decode seconds. One background-suspended clip is excluded from both engines. The first-use cost remains: Aidana’s first clip took 2.865 s versus 0.573 s for ONNX.ONNX · 2 CPU threads41.575ONNX · 2 CPU threads: 41.575 decode secondsAidana · MLX GPU12.400Aidana · MLX GPU: 12.400 decode seconds0decode seconds
  1. ONNX · 2 CPU threads41.575
  2. Aidana · MLX GPU12.400

decode seconds

One background-suspended clip is excluded from both engines. The first-use cost remains: Aidana’s first clip took 2.865 s versus 0.573 s for ONNX.

View data & source
The same 47 clips took less time to transcribe · decode seconds
ConfigurationValue
ONNX · 2 CPU threads41.575
Aidana · MLX GPU12.400

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)