Back to the experimentSupporting notebook

Qwen3.5-4B tmux status QLoRA

Source: lora/qwen35-4b-tmux-status-qlora-2026-09-24/MODEL_CARD.md · revision 6550ead3945b

This repository contains a PEFT LoRA adapter, a full merged BF16 Safetensors checkpoint in merged-bf16/, and a standalone AutoRound Q4_K_M GGUF for classifying a coding session’s latest 30 tmux terminal lines. It also includes tokenizer/template files and aggregate evaluation results. The adapter requires the pinned base Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a; the merged checkpoint and GGUF do not require a separate base download or LoRA merge. These artifacts are not native MLX models. The base model is Apache-2.0 licensed.

Output

The intended response is exactly one JSON object with a single state key: {"state":"running"}, {"state":"finished"}, or {"state":"needs_input"}. Use the included system_prompt.txt, a user message containing JSON {"terminal_last_30_lines":"..."}, and the Qwen3.5 chat template with enable_thinking=False. The included classifier.py implements deterministic constrained decoding for this contract. It does not represent calibrated class probabilities.

Training

Evaluation

Using the same deterministic three-label decoder on the 202-case held-out set:

Model Agreement with teacher Macro-F1 running recall finished recall needs_input recall
Base 97/202 (48.02%) 0.444 67/162 14/20 16/20
This LoRA 192/202 (95.05%) 0.863 160/162 16/20 16/20
Merged BF16 194/202 (96.04%) 0.887 161/162 16/20 17/20
AutoRound Q4_K_M GGUF 193/202 (95.54%) 0.879 160/162 16/20 17/20

Fine-tuned confusion matrix (rows: teacher labels, columns: predictions; order running, finished, needs_input): [[160,0,2],[2,16,2],[0,4,16]].

These are DeepSeek teacher predictions, not human-adjudicated truth. Only 20 examples represent each minority class; near-duplicate groups were split together, but sessions/projects can span splits. The test set was inspected during earlier experimentation. Raw terminal captures and per-record predictions are not published. The model has not been assessed on a new blind real-world set. A constrained decoder guarantees the JSON shape, not the accuracy of its choice.

Use

Use an environment with Transformers/PEFT support for Qwen3.5 (the training environment used Transformers 5.5.0, PEFT 0.21.0, Unsloth 2026.9.10). Load the base with 4-bit bitsandbytes and attach the adapter through PEFT. Match enable_thinking=False and system_prompt.txt. If you use unconstrained generation, validate the JSON response yourself. For the tested decoder see classifier.py.

For the standalone GGUF, load gguf/Qwen3.5-4B-Tmux-Status-AutoRound-HQ-Q4_K_M.gguf with a recent llama.cpp build that supports Qwen3.5 (b96806d9 was tested). The optional gguf/mmproj-model.gguf supplies the original unchanged vision pathway; this classifier was trained and evaluated on text only. Use OpenAI chat completions with chat_template_kwargs: {"enable_thinking": false} and a JSON schema restricting state to the three labels. The text GGUF SHA-256 is d2e90408505acdc276926bb537062127487f19bb87f1dbd9b842ae2b10bc53c1.

For the unquantized, already-merged model, download the merged-bf16/ folder and load that folder with a Qwen3.5-compatible Transformers version (5.5.0 was used for evaluation). It contains the three BF16 Safetensors shards, their index, model configuration, tokenizer, chat template, and unchanged vision pathway. Do not attach the LoRA adapter again. Use system_prompt.txt, enable_thinking=False, and the same constrained three-label decoder. This full checkpoint scored 194/202 (96.04%), macro-F1 0.887, on the held-out teacher-label set. The three shard SHA-256 hashes, in order, are 24979811f253650a6e83fe772c0236135280741b45ed9385ef451fbd40de05ac, 90fe1f6810842c2188537fbb26494e736d497d03b193141211f317efd597fe79, and bcacd4b6ced58e399d33a881370ab518048ec0bf265846490658f85065b8e24c.

The GGUF was produced by AutoRound 0.15.1 with 256 train-only calibration cases, 200 SignRound iterations, and the algorithm extension enabled. Its exporter recalculated packing scales on 248 of 258 quantized matrices after Qwen3.5 tensor transformations. The matrices being repacked came from AutoRound’s tuned dequantized weights, but this prevents a claim that every final GGUF scale was retained directly from optimization. The GGUF’s own 202-case evaluation above is the basis for its measured performance.

The adapter SHA-256 is 76b27f4823cfbd65959ba8d24fed3297ba38a5748e2bc0e80f27603ef780c27e. Full aggregate metrics, counts, validation curves, quantization metadata and deployment tests are in metrics/. The experiment report and training script are maintained in the private experimentos repository.