Back to the experimentSupporting notebook

Plan — Qwen3.8-Flash-Next EXL3 3.50 HQ + 3.75 MTP6

Source: qwen-next-flash/exl3-350hq-375-2026-09-15/PLAN.md · revision 6550ead3945b

Date: 2026-09-15 Host: laguna-s21 (3x RTX 3090 24 GB, 125 GB RAM, 608 GB free on /) Locked by user: proceed end-to-end; do not change GPU power limits.

This is the working plan. Status and evidence live in JOURNEY.md and data/. Runtime logs live in WORKSTATION/logs (too large to commit).

Goal

  1. Download official BF16 Qwen/Qwen3.8-Flash-Next (335.30 GiB, 131 safetensors).
  2. Quantize 3.50 bpw HQ with the user recipe (head 6 / ngram 6 / MTP 4 / vision 6).
  3. Upload that quant to groxaxo on Hugging Face.
  4. Serve it on proxy :12434 using the existing 4.05bpw EXL3 family params, with KV cache 8-bit, and measure:
    • decode tokens/s
    • max model length that actually loads/serves
  5. Quantize 3.75 bpw (head 6 / ngram 6 / MTP 6 / vision 6, HQ off) from the same BF16 tree.
  6. Upload 3.75 as well.
  7. Record the whole journey in this experiment folder.

Tooling (found on this machine)

Item Path / value
ExLlamaV3 repo WORKSTATION/exllamav3 (v1.5.0-1-g02aef45, editable install)
Convert entry WORKSTATION/convert.py
Python WORKSTATION/python (exllamav3 1.5.0+cu128.torch2.11.0)
HF account groxaxo (write token already in env)
Proxy qwen36-multi-proxy on 0.0.0.0:12434 (qwen36_variant_proxy.py)
4.05bpw EXL3 serve params family qwen38-flash-next-unc-exl3-4bpw in qwen36_model_families.json
4.05bpw reference weights WORKSTATION/Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 (101 G, 4.05 / h6 / mtp4, HQ off, vis not 6)
Unc 4.05 live backend WORKSTATION/Qwen3.8-Flash-Next-Unc-exl3-4bpw on :12703, CACHE_KV_BITS=6, CACHE_SIZE=180224, chunk 256, GPU split 23.5×3

Path aliases

The two user recipes disagree on prefix (/mnt/models vs /models). Both now resolve to the same tree:

Canonical locations:

Role Path
BF16 source /models/Qwen3.8-Flash-Next
3.50 HQ out /mnt/models/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6
3.50 work /mnt/models/exl3-work-qwen38-flash-350-hq
3.75 out /models/Qwen3.8-Flash-Next-exl3-3.75bpw_h6_ng6_mtp6
3.75 work /models/exl3-work-qwen38-flash-375-mtp6

Locked recipes

A — 3.50 bpw HQ (user script, unmodified flags)

bits=3.50  hq=ON  head=6  ngram=6  mtp=4  vision=6
out_scales=always  codebook=mul1  cal=250x2048  checkpoint_interval=120
devices=0,1,2  CUDA_VISIBLE_DEVICES=0,1,2

Derived from turboderp-style 4.05bpw_h6_ng6. Deltas vs that style: 4.05→3.50, HQ off→on. MTP stays 4. Vision is 6 (Unc 4.05 used vis 16).

B — 3.75 bpw MTP6 (user command, HQ off)

bits=3.75  hq=OFF  head=6  ngram=6  mtp=6  vision=6
out_scales=always  codebook=mul1  cal=250x2048
devices=0,1,2

Operational only (does not change bitrate): -cpi 120 so a crash can --resume.

GPU policy

Disk budget

Start: 608 G free. BF16 is 335 G.

Peak plan (sequential work dirs, delete work after each compile):

After Approx used from the 608 G Notes
BF16 download 335 G local-dir only; watch HF cache for duplication
3.50 convert 335 + ~90 work + ~90 out ≈ 515 G delete work on success
3.50 uploaded 335 + 90 = 425 G
3.75 convert 425 + ~95 work + ~95 out ≈ 615 G tight (~7 G)

Mitigations, in order, if free space < 40 G: drop Hugging Face hub cache (~46 G), delete the just-finished work dir, pause 3.75 until 3.50 work is gone. Never delete the BF16 tree until both quants compile.

Abort the wrapper if df available < 25 G.

Stages

  1. Download Qwen/Qwen3.8-Flash-Next → /models/Qwen3.8-Flash-Next with hf download --local-dir, HF_HUB_ENABLE_HF_TRANSFER=1, resumable. Success: config.json + model.safetensors.index.json + 131 *.safetensors.
  2. Convert 3.50 HQ via scripts/convert_350_hq.sh. --resume if work_dir/args.json exists.
  3. Upload 3.50 to groxaxo/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6 with a model card. Delete 3.50 work dir after a successful compile (keep out dir).
  4. Proxy + bench 3.50 on :12434:
    • Clone qwen38-flash-next-unc-exl3-4bpw serve env: CACHE_SIZE=180224, GPU_SPLIT=23.5,23.5,23.5, CHUNK_SIZE=256, MAX_BATCH_SIZE=1, NGRAM_RAM=1, Sharp template, layer-split, port 12705.
    • Override CACHE_KV_BITS=8 (4.05 live family uses 6; user asked 8-bit KV).
    • Restart proxy only when active_requests=0.
    • Measure short-stream tok/s (warmup + several 256-token generations).
    • Measure max model len: binary-search CACHE_SIZE that still loads with q8 KV, then a long prompt near that cap.
  5. Convert 3.75 MTP6 via scripts/convert_375_mtp6.sh, same BF16, --resume if needed.
  6. Upload 3.75 to groxaxo/Qwen3.8-Flash-Next-exl3-3.75bpw_h6_ng6_mtp6.
  7. Optional 3.75 serve/bench with the same 4.05-derived params + q8 KV (port 12706) after 3.50 numbers are recorded.
  8. Journal this folder (JOURNEY.md, data/*.json, README row in experimentos/README.md).

Watchdog

Quantization is multi-hour. Alongside the job:

Success criteria

Out of scope