Started: 2026-09-15. Status: 3.50 HQ live on :12434 (q8 KV 524288, gpu_split 24,24,24). 3.75 convert paused.
Timeline
| UTC-ish | What |
|---|---|
| 2026-09-15 session start | User asked to download official Flash-Next, quantize 3.50 HQ, upload, bench on :12434 with 4.05 params + 8-bit KV. Then: also run 3.75 MTP6, write the journey here, write the plan and remember it, proceed end-to-end. |
| 2026-09-15 | Located ExLlamaV3 at WORKSTATION/exllamav3 (v1.5.0-1-g02aef45). /models already aliases WORKSTATION/qwen. Created /mnt/models → same target so both recipes share one BF16 tree. 3x3090 idle, 245 W power left untouched, 608 G free, proxy :12434 idle. |
| 2026-09-15 | Plan written (PLAN.md), filed in MemPalace, download + convert watchdog launched. |
| 2026-09-15 | Pushed plan to groxaxo/experimentos c5bf147. BF16 download running to /models/Qwen3.8-Flash-Next with HF_XET_HIGH_PERFORMANCE=1. Power limits left at 245 W. |
| 2026-09-15 | claude -p --model haiku 429 (z.ai weekly limit until 2026-09-17 04:14). Swapped the 15-min cron to a bash Telegram status script so updates still fire. |
| 2026-09-15 11:18Z | BF16 download complete (336 G, 131 safetensors). |
| 2026-09-15 15:45Z | 3.50 HQ convert finished (-- All done). Work dir deleted. Disk output 93 G. Effective bits in quantization_config.json: 3.52 (HQ). |
| 2026-09-15 16:11Z | Uploaded https://huggingface.co/groxaxo/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6 (commit 7b55e0f). |
| 2026-09-15 16:11–18:43Z | 3.75 convert reached language layer 26 / next_module_idx=29, then the retry-wrapper was SIGTERM’d with no MONITOR_RESULT (GPUs later used for a brief q6 199936 serve that was shut down). |
| 2026-09-16 | q8 KV probe: 160000 tokens loads, 160256 and 180224 fail (Insufficient VRAM in split; cache must be a multiple of 256). 4.05 family’s 180224 q6 does not fit as q8. |
| 2026-09-16 | Direct :12705 bench, q8 / 160000 / chunk 256 / split 23.5×3: TTFT 0.236 s, 65.9 decode tok/s after TTFT (256 new tokens, length stop). VRAM 23248 / 22890 / 16604 MiB. Power limits still 245 W. |
| 2026-09-16 18:53Z | Retry-wrapper relaunched; 3.75 convert --resume from work dir. |
| 2026-09-16 | User: leave 3.75 aside; optimize 3.50; find q8 max-len sweet spot; put it on :12434 replacing Unc 4.05 EXL3. 3.75 paused (MONITOR_RESULT=PAUSED, work dir kept). |
| 2026-09-16 | q8 sweet spot confirmed 160000. 160256/163840/180224 fail. chunk 128 invalid (page size 256). GPU0 caps 20/19 did not unlock 163840. MTP left off (would steal VRAM from context). |
| 2026-09-16 | Disabled qwen38-flash-next-unc-exl3-4bpw. Added qwen38-flash-next-exl3-350bpw on :12703 with same public IDs (qwen38-unc-exl3-*). Proxy restart. Smoke: instruct → pong. Stream 256: 66.9 tok/s after TTFT. |
| 2026-09-16 | gpu_split 24,24,24 is not inert — 23.5 was the limiter. q8 load ceiling 606208 (622592 fail). q6 load ceiling 720896 (753664 fail). Live q8 moved to 524288 for decode headroom. Proxy stream 256: 67.9 tok/s. |
| 2026-09-16 | Concurrency 3: server dispatch now multiplexes up to 3 Jobs on one Generator (MAX_BATCH_SIZE=3). Live load batch=3 q8 524288 split 24,24,24. Three parallel 128-token requests through :12434 enqueued active=3/3; aggregate wall 17.3 s, 22 tok/s combined (per-stream slower than the 68 tok/s single-stream, as expected on 3-GPU layer-split). |
| 2026-09-16 | Standard moved to concurrency 2 + q8 400128 (~400k, page-aligned). Load OK, active=2/2. Dual 128-token: 2.7 s wall, ~47 tok/s each, 94.5 tok/s aggregate. VRAM free 979/1640/4206 MiB. |
| 2026-09-16 | Disabled remaining other EXL3 family (qwen38-27b-unc-exl3-8bpw). Proxy lists only qwen38-unc-exl3-*. FP16 KV at 400128 batch 2 loads and serves; dual 128-token 2.67 s, ~48 tok/s each, 95.8 agg. Left as default (CACHE_KV_BITS=0). GPU1 free 362 MiB. |
Decisions
- One BF16 copy. Recipe prefixes
/mnt/modelsand/modelsare aliases, not two downloads. - Convert with all three GPUs; serve with the 4.05 EXL3 layer-split layout. GPU-1-only is convert fallback only — 125B EXL3 will not fit on one 24 GB card.
- 4.05 live family uses
CACHE_KV_BITS=6/CACHE_SIZE=180224/ chunk 256. Bench clones that and sets KV to 8. - Delete each work dir after a successful compile so 3.75 can fit next to 3.50 + BF16.
- Do not touch nvidia power limits.
Running commands (as launched)
See scripts/. Runtime logs: WORKSTATION/logs.
Results (fill as stages complete)
- Download size / shard count: 336 G / 131 safetensors
- 3.50 out size /
quantization_config: 93 G, bits 3.52, head 6, mtp 4, vision 6, HQ recipe, cal 250×2048, mul1, out_scales always - 3.50 HF URL: https://huggingface.co/groxaxo/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6
- 3.50 tok/s / max model len (q8 KV): 65.9 decode tok/s after TTFT; max cache 160000 tokens (180224 q8 does not load). Evidence:
data/q8_kv_bench.json - 3.75 out size /
quantization_config: paused at module 29 (flags/paused375); work dir kept - 3.75 HF URL: pending
- Proxy: Unc 4.05 EXL3 family disabled. Official 3.50 HQ serves
:12434asqwen38-unc-exl3-*on:12703, q8 / 524288 / split 24,24,24. q8 ceiling 606208; q6 ceiling 720896. Proxy stream 67.9 tok/s.