Back to the experimentSupporting notebook

Release plan — Qwen3.5-4B tmux classifier

Source: lora/qwen35-4b-tmux-status-qlora-2026-09-24/PLAN.md · revision 6550ead3945b

Objective: publish the trained LoRA and aggregate findings to Hugging Face, merge the LoRA into the exact Qwen/Qwen3.5-4B revision, quantize merged weights locally on .54 using Intel AutoRound high-quality tuning, validate the quantized classifier against the frozen held-out set and the merged parent, then make it the tmux-pocket default behind an on-demand local service. Preserve fallback and unrelated GPU workloads.

  1. Record the existing reproducible LoRA run in the private experimentos/lora/... subfolder: exact base revision, adapter hash, dataset manifest/counts, full training configuration, all three validation losses, per-class base/fine-tuned metrics, coverage, error matrix, scripts, known limits. Add an index entry. Exclude raw terminal captures, prediction rows, credentials, and weights from Git.
  2. Use the Mac’s already-authenticated groxaxo Hugging Face identity to create and upload a model repository with the adapter, tokenizer/template, system prompt, model card and aggregate metrics. Verify remote file listing and adapter SHA-256. Do not upload private train/validation/test examples.
  3. On .54, preflight 21 GB disk free, 32 GB RAM/12 GB VRAM and existing GPU tenants; the exact pinned HF base snapshot is already present as 8.8 GB of symlinked blobs (du -shL), so do not redownload; use the existing training venv and preserve at least 3 GB disk slack. Merge LoRA into the exact BF16 base using the isolated Transformers/PEFT runtime, with vision frozen. Save merged BF16 to an isolated directory and verify finite weights, adapter effect and reload. Avoid deleting any previous artifact.
  4. Install AutoRound 0.15.1 in the existing isolated Transformers 5.5 training venv (without duplicating its 5.5 GB runtime), smoke-test its Qwen3.5 compatibility using the already cached base, then quantize the merged checkpoint using real optimized rounding (iters>0), representative train-only calibration, high quality 4-bit scheme with safe head/embedding treatment. Record package versions, command, seed, samples, format, layer config, output size. Keep quantization on .54. If resource/preemption conflicts arise, queue or choose low-memory mode; do not stop unrelated services.
  5. Evaluate deterministic three-label JSON decoding on all 202 held-out records for merged and quantized variants; compare agreement and macro-F1 against the 95.05% agreement / 0.8626 macro-F1 adapter reference; require merged agreement >=93%, quantized macro-F1 >=0.83, and no class recall falling below 0.70 and inspect failures. Run a cold-start and real HTTP API probe before deployment. Do not deploy a regression without clear evidence.
  6. Add dedicated on-demand local backend behind the established lazy gate pattern; preserve the fleet-managed disabled control, snapshot settings before change, and avoid preempting the canonical Nemotron service. Wire tmux-pocket’s provider to it, with DeepSeek as fallback. Preserve settings API behavior and other providers. Test classified terminal cases, lazy cold start/idle reap, fallbacks, and service restoration. Publish final quantized model and updated findings to Hugging Face through the Mac if write access is restored; the current groxaxo Mac token was verified read-only and returned HTTP 403 for repo creation. Commit/push the experiment report to the existing private Git remote.

Authority: user explicitly asked to publish, merge, AutoRound HQ quantize on .54, test, and make default lazy classifier. No third-party service interruptions authorized beyond routine safe backend starts. Existing test is teacher-labeled and historically inspected; new tests needed for independent quality estimate.