Date: 2026-09-15
Host: laguna-s21 (3x RTX 3090 24 GB, 125 GB RAM, 608 GB free on /)
Locked by user: proceed end-to-end; do not change GPU power limits.
This is the working plan. Status and evidence live in JOURNEY.md and
data/. Runtime logs live in WORKSTATION/logs
(too large to commit).
Goal
- Download official BF16
Qwen/Qwen3.8-Flash-Next(335.30 GiB, 131 safetensors). - Quantize 3.50 bpw HQ with the user recipe (head 6 / ngram 6 / MTP 4 / vision 6).
- Upload that quant to
groxaxoon Hugging Face. - Serve it on proxy
:12434using the existing 4.05bpw EXL3 family params, with KV cache 8-bit, and measure:- decode tokens/s
- max model length that actually loads/serves
- Quantize 3.75 bpw (head 6 / ngram 6 / MTP 6 / vision 6, HQ off) from the same BF16 tree.
- Upload 3.75 as well.
- Record the whole journey in this experiment folder.
Tooling (found on this machine)
| Item | Path / value |
|---|---|
| ExLlamaV3 repo | WORKSTATION/exllamav3 (v1.5.0-1-g02aef45, editable install) |
| Convert entry | WORKSTATION/convert.py |
| Python | WORKSTATION/python (exllamav3 1.5.0+cu128.torch2.11.0) |
| HF account | groxaxo (write token already in env) |
| Proxy | qwen36-multi-proxy on 0.0.0.0:12434 (qwen36_variant_proxy.py) |
| 4.05bpw EXL3 serve params | family qwen38-flash-next-unc-exl3-4bpw in qwen36_model_families.json |
| 4.05bpw reference weights | WORKSTATION/Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 (101 G, 4.05 / h6 / mtp4, HQ off, vis not 6) |
| Unc 4.05 live backend | WORKSTATION/Qwen3.8-Flash-Next-Unc-exl3-4bpw on :12703, CACHE_KV_BITS=6, CACHE_SIZE=180224, chunk 256, GPU split 23.5×3 |
Path aliases
The two user recipes disagree on prefix (/mnt/models vs /models). Both now resolve to the same tree:
/models→WORKSTATION/qwen(pre-existing)/mnt/models→WORKSTATION/qwen(created 2026-09-15;/mnt/modelsdid not exist)
Canonical locations:
| Role | Path |
|---|---|
| BF16 source | /models/Qwen3.8-Flash-Next |
| 3.50 HQ out | /mnt/models/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6 |
| 3.50 work | /mnt/models/exl3-work-qwen38-flash-350-hq |
| 3.75 out | /models/Qwen3.8-Flash-Next-exl3-3.75bpw_h6_ng6_mtp6 |
| 3.75 work | /models/exl3-work-qwen38-flash-375-mtp6 |
Locked recipes
A — 3.50 bpw HQ (user script, unmodified flags)
bits=3.50 hq=ON head=6 ngram=6 mtp=4 vision=6
out_scales=always codebook=mul1 cal=250x2048 checkpoint_interval=120
devices=0,1,2 CUDA_VISIBLE_DEVICES=0,1,2
Derived from turboderp-style 4.05bpw_h6_ng6. Deltas vs that style: 4.05→3.50, HQ off→on. MTP stays 4. Vision is 6 (Unc 4.05 used vis 16).
B — 3.75 bpw MTP6 (user command, HQ off)
bits=3.75 hq=OFF head=6 ngram=6 mtp=6 vision=6
out_scales=always codebook=mul1 cal=250x2048
devices=0,1,2
Operational only (does not change bitrate): -cpi 120 so a crash can --resume.
GPU policy
- Do not change power limits (cards currently 245 W).
- Conversion: GPUs 0,1,2.
convert.pyparallel mode is default; 3-way Hessian/quant helps. All three cards were idle at plan time. - Fallback: if multi-GPU convert fails in a way that is clearly device-count related, retry on GPU 1 only (
CUDA_VISIBLE_DEVICES=1 --devices 0inside that visible set). - Serving 125B A6B EXL3 cannot fit on GPU 1 alone (~90 GB weights). Proxy backends use the 4.05 family layout: layer-split autosplit on 0,1,2 (
supports_tp=Falsefor Qwen4Exp).
Disk budget
Start: 608 G free. BF16 is 335 G.
Peak plan (sequential work dirs, delete work after each compile):
| After | Approx used from the 608 G | Notes |
|---|---|---|
| BF16 download | 335 G | local-dir only; watch HF cache for duplication |
| 3.50 convert | 335 + ~90 work + ~90 out ≈ 515 G | delete work on success |
| 3.50 uploaded | 335 + 90 = 425 G | |
| 3.75 convert | 425 + ~95 work + ~95 out ≈ 615 G | tight (~7 G) |
Mitigations, in order, if free space < 40 G: drop Hugging Face hub cache (~46 G), delete the just-finished work dir, pause 3.75 until 3.50 work is gone. Never delete the BF16 tree until both quants compile.
Abort the wrapper if df available < 25 G.
Stages
- Download
Qwen/Qwen3.8-Flash-Next→/models/Qwen3.8-Flash-Nextwithhf download --local-dir,HF_HUB_ENABLE_HF_TRANSFER=1, resumable. Success:config.json+model.safetensors.index.json+ 131*.safetensors. - Convert 3.50 HQ via
scripts/convert_350_hq.sh.--resumeifwork_dir/args.jsonexists. - Upload 3.50 to
groxaxo/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6with a model card. Delete 3.50 work dir after a successful compile (keep out dir). - Proxy + bench 3.50 on
:12434:- Clone
qwen38-flash-next-unc-exl3-4bpwserve env:CACHE_SIZE=180224,GPU_SPLIT=23.5,23.5,23.5,CHUNK_SIZE=256,MAX_BATCH_SIZE=1,NGRAM_RAM=1, Sharp template, layer-split, port 12705. - Override
CACHE_KV_BITS=8(4.05 live family uses 6; user asked 8-bit KV). - Restart proxy only when
active_requests=0. - Measure short-stream tok/s (warmup + several 256-token generations).
- Measure max model len: binary-search
CACHE_SIZEthat still loads with q8 KV, then a long prompt near that cap.
- Clone
- Convert 3.75 MTP6 via
scripts/convert_375_mtp6.sh, same BF16,--resumeif needed. - Upload 3.75 to
groxaxo/Qwen3.8-Flash-Next-exl3-3.75bpw_h6_ng6_mtp6. - Optional 3.75 serve/bench with the same 4.05-derived params + q8 KV (port 12706) after 3.50 numbers are recorded.
- Journal this folder (
JOURNEY.md,data/*.json, README row inexperimentos/README.md).
Watchdog
Quantization is multi-hour. Alongside the job:
- Self-healing
scripts/retry-wrapper.sh(max 4 attempts, stage flags so a retry never re-downloads a finished tree). - Local crontab
*/15Haiku monitor → Telegram chat6025513237(no--dangerously-skip-permissions). - Grok
monitorthat stays silent untilMONITOR_RESULT=SUCCESS|GAVE_UP, disk < 25 G, or the wrapper dies.
Success criteria
- Both output dirs contain
quantization_config.json,model.safetensors.index.json,ngram_embedding.safetensors. - 3.50
quantization_config.bits == 3.50and HQ strategy present; 3.75bits == 3.75,mtp_bits == 6, no HQ. - Both Hugging Face repos exist and list the shards.
- 3.50 answers through
:12434with q8 KV. Recorded: tok/s, max cache tokens, VRAM per GPU, exact cmdline/env. - This plan + journey committed to
groxaxo/experimentos.
Out of scope
- Changing 3090 power limits.
- Replacing the live Unc 4.05
:12703backend. - Vision/mmproj bring-up (4.05 EXL3 family is text-only; same here unless load proves otherwise).
- Tensor parallel (not implemented for Qwen4Exp).