Journal

The EXL3 quantization journey

The notebook follows official BF16 weights through EXL3 conversion, artifact checks and deployment work.

Vlad / experimentos.
The boundary

The archive marks this experiment in progress; completion is not inferred from a plan or training process.

The question

The EXL3 work began with official BF16 weights and followed the conversion into a deployable artifact. The journey keeps the intermediate checks and remaining work visible rather than declaring success from a running process.

From the original notebook

Started: 2026-09-15. Status: 3.50 HQ live on :12434 (q8 KV 524288, gpu_split 24,24,24). 3.75 convert paused.

Timeline

UTC-ish What
2026-09-15 session start User asked to download official Flash-Next, quantize 3.50 HQ, upload, bench on :12434 with 4.05 params + 8-bit KV. Then: also run 3.75 MTP6, write the journey here, write the plan and remember it, proceed end-to-end.
2026-09-15 Located ExLlamaV3 at WORKSTATION/exllamav3 (v1.5.0-1-g02aef45). /models already aliases WORKSTATION/qwen. Created /mnt/models → same target so both recipes share one BF16 tree. 3x3090 idle, 245 W power left untouched, 608 G free, proxy :12434 idle.
2026-09-15 Plan written (PLAN.md), filed in MemPalace, download + convert watchdog launched.
2026-09-15 Pushed plan to groxaxo/experimentos c5bf147. BF16 download running to /models/Qwen3.8-Flash-Next with HF_XET_HIGH_PERFORMANCE=1. Power limits left at 245 W.
2026-09-15 claude -p --model haiku 429 (z.ai weekly limit until 2026-09-17 04:14). Swapped the 15-min cron to a bash Telegram status script so updates still fire.
2026-09-15 11:18Z BF16 download complete (336 G, 131 safetensors).
2026-09-15 15:45Z 3.50 HQ convert finished (-- All done). Work dir deleted. Disk output 93 G. Effective bits in quantization_config.json: 3.52 (HQ).
2026-09-15 16:11Z Uploaded https://huggingface.co/groxaxo/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6 (commit 7b55e0f).
2026-09-15 16:11–18:43Z 3.75 convert reached language layer 26 / next_module_idx=29, then the retry-wrapper was SIGTERM’d with no MONITOR_RESULT (GPUs later used for a brief q6 199936 serve that was shut down).
2026-09-16 q8 KV probe: 160000 tokens loads, 160256 and 180224 fail (Insufficient VRAM in split; cache must be a multiple of 256). 4.05 family’s 180224 q6 does not fit as q8.
2026-09-16 Direct :12705 bench, q8 / 160000 / chunk 256 / split 23.5×3: TTFT 0.236 s, 65.9 decode tok/s after TTFT (256 new tokens, length stop). VRAM 23248 / 22890 / 16604 MiB. Power limits still 245 W.
2026-09-16 18:53Z Retry-wrapper relaunched; 3.75 convert --resume from work dir.
2026-09-16 User: leave 3.75 aside; optimize 3.50; find q8 max-len sweet spot; put it on :12434 replacing Unc 4.05 EXL3. 3.75 paused (MONITOR_RESULT=PAUSED, work dir kept).
2026-09-16 q8 sweet spot confirmed 160000. 160256/163840/180224 fail. chunk 128 invalid (page size 256). GPU0 caps 20/19 did not unlock 163840. MTP left off (would steal VRAM from context).
2026-09-16 Disabled qwen38-flash-next-unc-exl3-4bpw. Added qwen38-flash-next-exl3-350bpw on :12703 with same public IDs (qwen38-unc-exl3-*). Proxy restart. Smoke: instruct → pong. Stream 256: 66.9 tok/s after TTFT.
2026-09-16 gpu_split 24,24,24 is not inert — 23.5 was the limiter. q8 load ceiling 606208 (622592 fail). q6 load ceiling 720896 (753664 fail). Live q8 moved to 524288 for decode headroom. Proxy stream 256: 67.9 tok/s.
2026-09-16 Concurrency 3: server dispatch now multiplexes up to 3 Jobs on one Generator (MAX_BATCH_SIZE=3). Live load batch=3 q8 524288 split 24,24,24. Three parallel 128-token requests through :12434 enqueued active=3/3; aggregate wall 17.3 s, 22 tok/s combined (per-stream slower than the 68 tok/s single-stream, as expected on 3-GPU layer-split).
2026-09-16 Standard moved to concurrency 2 + q8 400128 (~400k, page-aligned). Load OK, active=2/2. Dual 128-token: 2.7 s wall, ~47 tok/s each, 94.5 tok/s aggregate. VRAM free 979/1640/4206 MiB.
2026-09-16 Disabled remaining other EXL3 family (qwen38-27b-unc-exl3-8bpw). Proxy lists only qwen38-unc-exl3-*. FP16 KV at 400128 batch 2 loads and serves; dual 128-token 2.67 s, ~48 tok/s each, 95.8 agg. Left as default (CACHE_KV_BITS=0). GPU1 free 362 MiB.

Decisions

  • One BF16 copy. Recipe prefixes /mnt/models and /models are aliases, not two downloads.
  • Convert with all three GPUs; serve with the 4.05 EXL3 layer-split layout. GPU-1-only is convert fallback only — 125B EXL3 will not fit on one 24 GB card.
  • 4.05 live family uses CACHE_KV_BITS=6 / CACHE_SIZE=180224 / chunk 256. Bench clones that and sets KV to 8.
  • Delete each work dir after a successful compile so 3.75 can fit next to 3.50 + BF16.
  • Do not touch nvidia power limits.

Running commands (as launched)

See scripts/. Runtime logs: WORKSTATION/logs.

Results (fill as stages complete)

  • Download size / shard count: 336 G / 131 safetensors
  • 3.50 out size / quantization_config: 93 G, bits 3.52, head 6, mtp 4, vision 6, HQ recipe, cal 250×2048, mul1, out_scales always
  • 3.50 HF URL: https://huggingface.co/groxaxo/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6
  • 3.50 tok/s / max model len (q8 KV): 65.9 decode tok/s after TTFT; max cache 160000 tokens (180224 q8 does not load). Evidence: data/q8_kv_bench.json
  • 3.75 out size / quantization_config: paused at module 29 (flags/paused375); work dir kept
  • 3.75 HF URL: pending
  • Proxy: Unc 4.05 EXL3 family disabled. Official 3.50 HQ serves :12434 as qwen38-unc-exl3-* on :12703, q8 / 524288 / split 24,24,24. q8 ceiling 606208; q6 ceiling 720896. Proxy stream 67.9 tok/s.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (3)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS