Full-spec HQ4 build (256 samples / 200 iters) killed at block 6/36 (~17%) — did NOT complete. No traceback, no reboot, no OOM signature: wrapper + quant job vanished together between 12:47 and 14:01 NZST. Most probable cause: external SIGKILL/SIGTERM from parallel session work (killer unidentified). Not relaunched — paused per user instruction.
See JOURNEY.md for the forensic timeline. Machine-readable facts under docs/.
Evidence
| Signal | Value |
|---|---|
| Last progress | layers.6: 6/36, 12:47:49 NZST, ~280 s/block, block loss converging (115→9.4) |
| Wrapper diag log | Ends at 12:19 launch line — no exit code, no diagnostics, no MONITOR_RESULT |
| Reboot? | No (uptime continuous, 2 days 21:45) |
| Kernel GPU fault (Xid)? | None in 12:30–14:05 window (journal readable) |
| OOM killer (kernel + userspace)? | No signature in readable logs |
| Monitor | Detected case-4 (wrapper dead, retries unexhausted), Telegram’d, kept cron |
| Partial outputs | work/hq4-full/ holds only backbone_stub/ — 6 blocks of progress fully lost (AutoRound has no resume) |
Lessons (applied to future runs)
pkill -f quantizematchestools/quantize_moss_autoround_awq.py— any peer cleanup for their quant work can kill ours. Use pidfiles + narrow patterns on both sides.- Wrappers should
trap SIGTERMand log it — narrows future forensics (SIGKILL stays invisible). pgrep -f <name>self-matches the monitoring agent’s own cmdline (prompt contains the name) → false “alive”. Usepgrep -f "[n]ame"form plus pidfile checks.- Long AutoRound runs are all-or-nothing — budget for restarts, keep attempts ≥ 2.