Back to the experimentSupporting notebook

Journey — Qwen3.8-Flash-Next EXL3 3.50 HQ + 3.75 MTP6

Source: qwen-next-flash/exl3-350hq-375-2026-09-15/JOURNEY.md · revision 6550ead3945b

Started: 2026-09-15. Status: 3.50 HQ live on :12434 (q8 KV 524288, gpu_split 24,24,24). 3.75 convert paused.

Timeline

UTC-ish What
2026-09-15 session start User asked to download official Flash-Next, quantize 3.50 HQ, upload, bench on :12434 with 4.05 params + 8-bit KV. Then: also run 3.75 MTP6, write the journey here, write the plan and remember it, proceed end-to-end.
2026-09-15 Located ExLlamaV3 at WORKSTATION/exllamav3 (v1.5.0-1-g02aef45). /models already aliases WORKSTATION/qwen. Created /mnt/models → same target so both recipes share one BF16 tree. 3x3090 idle, 245 W power left untouched, 608 G free, proxy :12434 idle.
2026-09-15 Plan written (PLAN.md), filed in MemPalace, download + convert watchdog launched.
2026-09-15 Pushed plan to groxaxo/experimentos c5bf147. BF16 download running to /models/Qwen3.8-Flash-Next with HF_XET_HIGH_PERFORMANCE=1. Power limits left at 245 W.
2026-09-15 claude -p --model haiku 429 (z.ai weekly limit until 2026-09-17 04:14). Swapped the 15-min cron to a bash Telegram status script so updates still fire.
2026-09-15 11:18Z BF16 download complete (336 G, 131 safetensors).
2026-09-15 15:45Z 3.50 HQ convert finished (-- All done). Work dir deleted. Disk output 93 G. Effective bits in quantization_config.json: 3.52 (HQ).
2026-09-15 16:11Z Uploaded https://huggingface.co/groxaxo/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6 (commit 7b55e0f).
2026-09-15 16:11–18:43Z 3.75 convert reached language layer 26 / next_module_idx=29, then the retry-wrapper was SIGTERM’d with no MONITOR_RESULT (GPUs later used for a brief q6 199936 serve that was shut down).
2026-09-16 q8 KV probe: 160000 tokens loads, 160256 and 180224 fail (Insufficient VRAM in split; cache must be a multiple of 256). 4.05 family’s 180224 q6 does not fit as q8.
2026-09-16 Direct :12705 bench, q8 / 160000 / chunk 256 / split 23.5×3: TTFT 0.236 s, 65.9 decode tok/s after TTFT (256 new tokens, length stop). VRAM 23248 / 22890 / 16604 MiB. Power limits still 245 W.
2026-09-16 18:53Z Retry-wrapper relaunched; 3.75 convert --resume from work dir.
2026-09-16 User: leave 3.75 aside; optimize 3.50; find q8 max-len sweet spot; put it on :12434 replacing Unc 4.05 EXL3. 3.75 paused (MONITOR_RESULT=PAUSED, work dir kept).
2026-09-16 q8 sweet spot confirmed 160000. 160256/163840/180224 fail. chunk 128 invalid (page size 256). GPU0 caps 20/19 did not unlock 163840. MTP left off (would steal VRAM from context).
2026-09-16 Disabled qwen38-flash-next-unc-exl3-4bpw. Added qwen38-flash-next-exl3-350bpw on :12703 with same public IDs (qwen38-unc-exl3-*). Proxy restart. Smoke: instruct → pong. Stream 256: 66.9 tok/s after TTFT.
2026-09-16 gpu_split 24,24,24 is not inert — 23.5 was the limiter. q8 load ceiling 606208 (622592 fail). q6 load ceiling 720896 (753664 fail). Live q8 moved to 524288 for decode headroom. Proxy stream 256: 67.9 tok/s.
2026-09-16 Concurrency 3: server dispatch now multiplexes up to 3 Jobs on one Generator (MAX_BATCH_SIZE=3). Live load batch=3 q8 524288 split 24,24,24. Three parallel 128-token requests through :12434 enqueued active=3/3; aggregate wall 17.3 s, 22 tok/s combined (per-stream slower than the 68 tok/s single-stream, as expected on 3-GPU layer-split).
2026-09-16 Standard moved to concurrency 2 + q8 400128 (~400k, page-aligned). Load OK, active=2/2. Dual 128-token: 2.7 s wall, ~47 tok/s each, 94.5 tok/s aggregate. VRAM free 979/1640/4206 MiB.
2026-09-16 Disabled remaining other EXL3 family (qwen38-27b-unc-exl3-8bpw). Proxy lists only qwen38-unc-exl3-*. FP16 KV at 400128 batch 2 loads and serves; dual 128-token 2.67 s, ~48 tok/s each, 95.8 agg. Left as default (CACHE_KV_BITS=0). GPU1 free 362 MiB.

Decisions

Running commands (as launched)

See scripts/. Runtime logs: WORKSTATION/logs.

Results (fill as stages complete)