The question
The tmux status problem is narrow enough to ask whether a small trained classifier can handle it locally. The project followed teacher-labeled examples through QLoRA, merged weights, quantization and a lazy service, while retaining the limits of that teacher agreement.
From the original notebook
24 September 2026. A text-only LoRA adapter for the three labels running, finished, and needs_input. The base is Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. Training and inference use the Qwen3.5 chat template with enable_thinking=False. The adapter is separate from the base and is not an MLX model. Repository: groxaxo/Qwen3.5-4B-Tmux-Status-QLoRA (release pending verification).
Purpose and output contract
Input: the coding session’s latest 30 tmux terminal lines in a JSON user message under terminal_last_30_lines. Use docs/system_prompt.txt. The response must be one of these exact JSON objects:
{"state":"running"}
{"state":"finished"}
{"state":"needs_input"}
The decoder constrains generation to those labels. That guarantees syntax and the label set, not correctness. The actual tmux product currently has an additional unknown state; the trained model does not predict it. Timeouts/provider failures can still use unknown operationally.
Source and split
| Set | running | finished | needs_input | Total |
|---|---|---|---|---|
| Train | 1,300 | 111 | 72 | 1,483 |
| Validation | 162 | 20 | 20 | 202 |
| Test | 162 | 20 | 20 | 202 |
The source had 3,839 valid DeepSeek-labeled records, reduced to 1,887 usable unique examples after request deduplication, normalization, conflict removal, and near-duplicate grouping. Six conflicting text buckets were excluded. Near-duplicate groups were kept within splits (max cross-split char n-gram cosine 0.875). Sessions and projects may still span splits. Labels were produced by DeepSeek and were not human adjudicated. The test set was inspected in prior work, so it is a fixed comparison set, not a new blind benchmark. Raw terminal captures remain private on the workstation.
The manifest contains source counts and snapshot/prompt hashes. All 1,483 train IDs were observed exactly once in each of three epochs, with no missing IDs and no duplicates within an epoch; see coverage. No example exceeded the 2,048-token limit.
Training
| Parameter | Value |
|---|---|
| Base | Qwen/Qwen3.5-4B (pinned revision above) |
| Method | QLoRA; 4-bit base; 42,467,328 trainable parameters (0.93%) |
| LoRA | rank 32, alpha 64, dropout 0; attention and MLP projections |
| GPU | one NVIDIA RTX 3060, 12 GB |
| Precision | BF16 compute |
| Batch | microbatch 1, gradient accumulation 8 |
| Peak LR | 2e-4; cosine decay; 5% warmup |
| Optimizer | AdamW 8-bit; weight decay 0.01; gradient clip 1.0 |
| Epochs | 3; lowest validation loss checkpoint selected |
| Loss | assistant answer tokens only; prompt/system/user masked |
| Sampling | shuffle without replacement; complete coverage audited |
| Training time | 4,603 seconds (~77 minutes) |
The exact script, configuration, and package versions are preserved. The run used local Trackio logging and never pushed raw data to the Hub.
Measured results
| Checkpoint | Validation loss (202) |
|---|---|
| Epoch 1, step 186 | 0.041414 |
| Epoch 2, step 372 | 0.031760 |
| Epoch 3, step 558, selected | 0.030031 |
On the same 202-case test split, with the same constrained JSON decoder:
| Model | Correct | Accuracy / teacher agreement | Macro-F1 |
|---|---|---|---|
| Base Qwen3.5-4B | 97/202 | 48.02% | 0.444 |
| Trained LoRA | 192/202 | 95.05% | 0.863 |
LoRA per-class recall: running 160/162 (98.77%), finished 16/20 (80%), needs_input 16/20 (80%). Confusion matrix (true rows, predicted columns; order running, finished, needs_input):
[[160, 0, 2],
[ 2,16, 2],
[ 0, 4,16]]
The ten mismatches were two running→needs_input, two finished→running, two finished→needs_input, and four needs_input→finished. Their raw text and request IDs are omitted from public material. The base and fine-tuned metric files preserve the full aggregate reports.
Interpretation and limitations
The 95.05% figure is agreement with automated teacher labels, not human-verified accuracy. The test has only 20 examples each of the two minority classes and may contain session/project overlap with train. Constrained decoding guarantees format but does not ensure the chosen label is right. It is inappropriate to infer broad generalization or claim a 47-point real-world gain from this test alone. The practical next test is a human-reviewed, fresh tmux challenge set with special attention to finished versus needs_input.
The LoRA was trained and tested under 4-bit base loading. An AutoRound quantized, merged model must be tested separately before it becomes the default. No merged or quantized accuracy results are attributed to this checkpoint yet.
Merge and AutoRound release work
The adapter was merged into the pinned BF16 base on research workstation by applying its 128 LoRA deltas to the corresponding text weights. The 738 serialized weight keys exactly match the original base, including 15 unchanged MTP tensors. The vision encoder remains unchanged. Transformers reloaded the merged checkpoint with zero missing, unexpected, mismatched, or error keys. The merged BF16 model was then evaluated on all 202 held-out cases with the same constrained decoder: 194/202 (96.04%), macro-F1 0.887; per-class recall running 161/162, finished 16/20, needs_input 17/20. See merge validation and the reproducible merge, normalization, and BF16 evaluator scripts.
AutoRound 0.15.1 high-quality GGUF Q4_K_M tuning completed on .54 with 256 train-only calibration examples, 200 iterations, and its algorithm extension enabled. The calibration set contains 128 running, 64 finished, and 64 needs_input examples. It is private and excluded from this repository. The quantization script records the exact inputs and settings. During quantization and the real server test, the existing Nemotron embedding service remained active.
The 2.82 GB text GGUF loaded in llama.cpp b96806d9 and completed the full frozen 202-case GGUF evaluation with constrained JSON schema decoding. It scored 193/202 (95.54%), macro-F1 0.879, and per-class recall running 160/162, finished 16/20, needs_input 17/20. It agreed with the LoRA adapter on 201/202 cases and the merged BF16 model on 201/202 cases. The GGUF changed one needs_input case in the correct direction compared with the adapter, and one running case in the wrong direction compared with merged BF16. These are teacher-label agreements, subject to the test limits above. Exact metrics, GGUF tensor types, sizes and hashes are in quantization validation.
The exporter recalculated quantization parameters for 248/258 quantized matrices after Qwen3.5 weight transformations. AutoRound first tuned and copied its dequantized weights into the live model, so those tuned weights were the inputs to GGUF packing, but the final packing scales for these matrices were recalculated. This limits any claim that the final GGUF retained every optimized quantization parameter. Its measured classifier result is the release evidence.
Deployment on .54
The text GGUF is now the tmux-pocket default through a loopback-only lazy gate at localhost; the llama.cpp GPU backend starts on demand at :18174 and stops after 300 seconds idle. A separate loopback-only lazy CPU backend at :18175/:18176 is the first fallback if GPU inference fails. DeepSeek remains the tertiary fallback, but a live probe returned HTTP 402, so it is not currently usable. Neither gate preempts the Nemotron embedding service. The four unit definitions are in deploy, and the cold-start, idle, GPU, CPU and fallback probes are in deployment validation. The tmux monitor’s existing disabled control was preserved; its saved provider setting is the local model.
The public Hugging Face release was published on 2026-09-25 with the LoRA adapter, standalone AutoRound GGUF, optional projector, tokenizer/template, model card, and aggregate findings. The publication script verified repository write access, local artifact hashes, the remote file listing, and the remote LFS SHA-256 hashes of all three weight files. A scoped groxaxo credential stored on .51 was installed on the Mac and this .54 host without exposing the token in logs; .54 performed the final upload because the Mac transfer was unstable. The aggregate report is also committed in the private experimentos Git repository.
On 2026-09-25 the existing, evaluated full BF16 merge was added under merged-bf16/. This is a complete Transformers checkpoint with the LoRA delta already applied; no separate adapter is needed. Its three Safetensors shards contain 738 BF16 tensors and total 9,319,820,640 weight-data bytes. The local index and all shard headers were checked, and Hugging Face LFS metadata returned the exact local SHA-256 for each shard. The root model card now explains how to use the merged checkpoint. The release retained the adapter and GGUF for users who want those formats.
Provenance
- Adapter SHA-256:
76b27f4823cfbd65959ba8d24fed3297ba38a5748e2bc0e80f27603ef780c27e. - Merged BF16 shard SHA-256 (00001, 00002, 00003):
24979811f253650a6e83fe772c0236135280741b45ed9385ef451fbd40de05ac,90fe1f6810842c2188537fbb26494e736d497d03b193141211f317efd597fe79,bcacd4b6ced58e399d33a881370ab518048ec0bf265846490658f85065b8e24c. - Original training artifact:
WORKSTATION/adapteronresearch workstation. - Frozen dataset snapshot SHA-256:
e37cde0573bad7d84681d60c57254c3dfba7df3fc6fc06ab590fbe99efb4ec14. - No dataset rows, terminal captures, or credentials are committed to this experiment or published with the model.
Keep exploring
- MODEL CARD.md — qwen35-4b-tmux-status-qlora-2026-09-24
- PLAN.md — qwen35-4b-tmux-status-qlora-2026-09-24
- PLAN REVIEW.md — qwen35-4b-tmux-status-qlora-2026-09-24
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.