Objective: publish the trained LoRA and aggregate findings to Hugging Face, merge the LoRA into the exact Qwen/Qwen3.5-4B revision, quantize merged weights locally on .54 using Intel AutoRound high-quality tuning, validate the quantized classifier against the frozen held-out set and the merged parent, then make it the tmux-pocket default behind an on-demand local service. Preserve fallback and unrelated GPU workloads.
- Record the existing reproducible LoRA run in the private
experimentos/lora/...subfolder: exact base revision, adapter hash, dataset manifest/counts, full training configuration, all three validation losses, per-class base/fine-tuned metrics, coverage, error matrix, scripts, known limits. Add an index entry. Exclude raw terminal captures, prediction rows, credentials, and weights from Git. - Use the Mac’s already-authenticated
groxaxoHugging Face identity to create and upload a model repository with the adapter, tokenizer/template, system prompt, model card and aggregate metrics. Verify remote file listing and adapter SHA-256. Do not upload private train/validation/test examples. - On .54, preflight 21 GB disk free, 32 GB RAM/12 GB VRAM and existing GPU tenants; the exact pinned HF base snapshot is already present as 8.8 GB of symlinked blobs (
du -shL), so do not redownload; use the existing training venv and preserve at least 3 GB disk slack. Merge LoRA into the exact BF16 base using the isolated Transformers/PEFT runtime, with vision frozen. Save merged BF16 to an isolated directory and verify finite weights, adapter effect and reload. Avoid deleting any previous artifact. - Install AutoRound 0.15.1 in the existing isolated Transformers 5.5 training venv (without duplicating its 5.5 GB runtime), smoke-test its Qwen3.5 compatibility using the already cached base, then quantize the merged checkpoint using real optimized rounding (
iters>0), representative train-only calibration, high quality 4-bit scheme with safe head/embedding treatment. Record package versions, command, seed, samples, format, layer config, output size. Keep quantization on .54. If resource/preemption conflicts arise, queue or choose low-memory mode; do not stop unrelated services. - Evaluate deterministic three-label JSON decoding on all 202 held-out records for merged and quantized variants; compare agreement and macro-F1 against the 95.05% agreement / 0.8626 macro-F1 adapter reference; require merged agreement >=93%, quantized macro-F1 >=0.83, and no class recall falling below 0.70 and inspect failures. Run a cold-start and real HTTP API probe before deployment. Do not deploy a regression without clear evidence.
- Add dedicated on-demand local backend behind the established lazy gate pattern; preserve the fleet-managed disabled control, snapshot settings before change, and avoid preempting the canonical Nemotron service. Wire tmux-pocket’s provider to it, with DeepSeek as fallback. Preserve settings API behavior and other providers. Test classified terminal cases, lazy cold start/idle reap, fallbacks, and service restoration. Publish final quantized model and updated findings to Hugging Face through the Mac if write access is restored; the current
groxaxoMac token was verified read-only and returned HTTP 403 for repo creation. Commit/push the experiment report to the existing private Git remote.
Authority: user explicitly asked to publish, merge, AutoRound HQ quantize on .54, test, and make default lazy classifier. No third-party service interruptions authorized beyond routine safe backend starts. Existing test is teacher-labeled and historically inspected; new tests needed for independent quality estimate.