Back to the experimentSupporting notebook

Plan Review — Qwen3.5-4B LoRA × AutoRound HQ 4-bit merge/quantize/deploy

Source: lora/qwen35-4b-tmux-status-qlora-2026-09-24/PLAN_REVIEW.md · revision 6550ead3945b

Now let me examine the training artifacts location, GPU tenants, HF cache, and the lazy-service pattern the plan references. One final check to tighten the disk-budget estimate (vision tower size) and confirm the experiment README absence:

1. Verdict — REVISE

The plan is grounded in real, verified local state (trained adapter, dataset, lazy-gate pattern, AutoRound arch support all exist), but two material gaps must be closed before execution: the disk budget only closes with an explicit venv/reclaim strategy the plan currently forbids-by-wording, and the only locally installed AutoRound cannot load qwen3_5, so the “install isolated AutoRound compatible with Qwen3.5” step is an unproven compatibility pairing, not a routine install.

2. Risks (prioritized)

  1. Disk exhaustion (operational, HIGH). Root fs is 96% full with 21 GB free. Budget: base BF16 download ≈7.5–8 GB (weights not in HF cache — only 9.9 MB of config) + merged BF16 save ≈7.5–8 GB + quantized output ≈2.5–3 GB ≈ 18–19 GB before any venv or pip cache. A fresh AutoRound venv (~5 GB; the training venv is 5.5 GB) does not fit. PLAN.md line 7 says “Avoid deleting any previous artifact”, which blocks the only reclaim path (deleting the re-downloadable base cache blobs after merge verification).
  2. AutoRound ↔ Qwen3.5 compatibility unproven (validation, HIGH). The sole local AutoRound (0.15.1 in WORKSTATION/.venv) ships Qwen3_5ForConditionalGeneration mappings (good) but runs transformers 5.0.0, which lacks models/qwen3_5 and cannot load the model; the arch exists only in transformers 5.5.0 (training venv). Pairing auto-round 0.15.1 with transformers 5.5.0 is plausible (Requires-Dist: transformers>=4.38, open-ended) but untested anywhere on this host.
  3. VRAM contention with a canonical tenant (safety, MEDIUM). 12,288 MiB total with 2,557 MiB held by the MemPalace canonical nemotron-embed vLLM (PID 3850034; gpu-memory-utilization 0.38) that must not be stopped. ~9.7 GiB free makes HQ iters>0 tuning of a 4B VL model plausible but tight (bf16 vision tower + activations), and other lazy GPU services can wake concurrently. Merged-BF16 eval (step 5) needs ~8 GB — GPU is borderline; CPU fallback should be stated.
  4. Fallback wiring is new, unspecified logic (scope, MEDIUM). session_status.py has a single provider and no fallback chain: on error it sets state=unknown. “Local default with DeepSeek fallback” is a real code change in classify()/poll() that must preserve: HTTPS-only-remote/HTTP-local validation (loopback endpoint is allowed), credential_source=opencode restricted to api.deepseek.com, settings-revision semantics, and the fleet_managed coordinated-poll path. The serving stack for the AutoRound int4 output is also unnamed (vLLM 0.25.0 exists; qwen3_5 int4 support in it unverified; transformers-serving is the fallback).
  5. Deployment driver mismatch (sequencing, LOW-MED). tmux-pocket-status.timer is not-found/failed (not installed in $WORKSTATION/.config/systemd/user); current operation is via coordinated polls (control.json: fleet_managed=true, 184 requests today). “Service restoration” tests must target the actual driver (hermes-coordinated path), not only the timer.
  6. Regression gate conflated (validation, LOW). PLAN.md line 9 compares macro-F1 against “the 95.0% adapter reference” — 95.05% is teacher agreement; macro-F1 is 0.8626. Minority classes have n=20 (high variance). No numeric accept/reject threshold is pre-declared, and the merged-BF16-vs-adapter delta (QLoRA trained on 4-bit base, merged into BF16) is a known shift source the gate should cover explicitly.
  7. Rollback gap (LOW). No stated snapshot/restore of settings.json/control.json before making the local backend default, and step 6’s “publish final quantized model to HF” from .54 has no stated credential path (step 2 only authorizes the Mac’s groxaxo identity).

3. Evidence (all directly observed unless noted)

Finding Evidence
Plan steps/wording PLAN.md lines 3–12 (line 7 preflight/“Avoid deleting”; line 8 AutoRound; line 9 “95.0% adapter reference”; line 10 fallback/timer/publish)
Adapter/run verified WORKSTATION/adapter_model.safetensors SHA-256 76b27f48… matches docs/summary.json; status.json phase=complete; dataset/train.jsonl present for train-only calibration; adapter_config.json pins base revision 851bf6e…
Reference metrics docs/summary.json + docs/test_finetuned_metrics.json: agreement 0.9505, macro-F1 0.86255, classes n=162/20/20
Base not cached; model is VL du: $WORKSTATION/.cache/huggingface/hub/models--Qwen--Qwen3.5-4B = 9.9 MB; cached config.json: Qwen3_5ForConditionalGeneration, vision_config present, text 32L×2560h, vocab 248,320 → est. 4.0B params ⇒ BF16 ≈ 7.5 GiB (inference from config, ±15%)
Disk/RAM/VRAM df -h / → 21 GB avail, 96% used; free -g → 19 GB available + 11 GB swap; nvidia-smi → 12,288 MiB total, 2,557 MiB used by VLLM::EngineCore = nemotron-embed.service (systemctl --user cat shows vLLM serve Nemotron AWQ, port 8001, gate :8000)
AutoRound state WORKSTATION/auto_round-0.15.1.dist-info + algorithms/transforms/awq/mappings.py:425 (Qwen3_5ForConditionalGeneration); MOSS venv transformers 5.0.0 without models/qwen3_5 vs WORKSTATION/tmux-qwen35-train transformers 5.5.0 with it; METADATA transformers>=4.38
Lazy-gate pattern WORKSTATION/lazy_systemd_http_proxy.py (on-demand start, IDLE_SECONDS reap, PREEMPT_SERVICES); $WORKSTATION/.config/systemd/user/qwen35-defiant-lazy-gate.service, chatterbox-lazy-proxy.service
tmux-pocket wiring constraints WORKSTATION/session_status.py lines 24–31 (single DeepSeek default), 145–156 (local-HTTP-only + opencode key locked to DeepSeek), 347–398 (no fallback chain), 493–495 (error→unknown), 401–421 (fleet gate); $WORKSTATION/.local/state/tmux-pocket/control.json fleet_managed:true; timer not-found/failed per systemctl --user status
Repo/index state WORKSTATION/experimentos git main→github.com/groxaxo/experimentos.git, lora/ untracked; README.md:13 index entry already present
Concurrent active work README.md (22:47:55) and MODEL_CARD.md (22:48:23) were created during this review (initial listing 22:47 showed neither); README names target repo groxaxo/Qwen3.5-4B-Tmux-Status-QLoRA — step 1 is being executed in parallel

4. Improved plan (minimal revision)

  1. Reorder step 4 before step 3’s download: first hardlink-clone tmux-qwen35-train (≈0 disk) + pip install auto-round, then run a smoke test (8 train samples, iters=1) proving the qwen3_5 + auto-round pairing loads and exports. Abort/escalate if it fails — before spending ~16 GB of the 21 GB budget.
  2. Add a disk budget gate and reclaim step: preflight df must show ≥ (base ≈8 GB + merged ≈8 GB + quant ≈3 GB + 2 GB slack); after merged-checkpoint verification passes, explicitly delete the downloaded base revision blobs from the HF cache (re-downloadable, not a previous artifact — amend the line-7 wording to permit exactly this).
  3. Pin the quantization runtime window: GPU ceiling via existing patterns, PYTORCH_CUDA_ALLOC_CONF, no preemption of nemotron-embed.service, and state that merged-BF16 eval runs on CPU if VRAM < ~9 GB free.
  4. Specify serving + fallback semantics: local OpenAI-compatible backend behind a new lazy gate (own user units, no PREEMPT of canonical services); poll tries local first, falls back to the existing DeepSeek provider on network/timeout/invalid, records which provider answered; settings snapshot (settings.json+control.json) taken before wiring and restorable to roll back; restoration test covers the coordinated-poll driver, not just the timer.
  5. Pre-declare the numeric gate before running step 5, e.g.: merged-BF16 within 2 pts agreement of the 95.05% adapter reference; quantized macro-F1 ≥ 0.83 (ref 0.8626); no per-class recall drop > 0.10; any violation ⇒ no deploy.
  6. Name the .54→HF credential path for the step-6 quantized publish (transfer via the Mac as in step 2, or a redacted token flow that never lands in project files).

5. Validation performed (read-only)

6. Changed files and external effects

None. No files created/edited/deleted; no installs, commits, pushes, service restarts, messages, or HF operations. (Note: another process created README.md/MODEL_CARD.md in the evidence root during the review — preserved untouched.)

7. Uncertainty