Back to the experimentSupporting notebook

moss-tts-hq4 / phase1-incident-2026-09

Source: moss-tts-hq4/phase1-incident-2026-09/README.md · revision 6550ead3945b

Full-spec HQ4 build (256 samples / 200 iters) killed at block 6/36 (~17%) — did NOT complete. No traceback, no reboot, no OOM signature: wrapper + quant job vanished together between 12:47 and 14:01 NZST. Most probable cause: external SIGKILL/SIGTERM from parallel session work (killer unidentified). Not relaunched — paused per user instruction.

See JOURNEY.md for the forensic timeline. Machine-readable facts under docs/.

Evidence

Signal Value
Last progress layers.6: 6/36, 12:47:49 NZST, ~280 s/block, block loss converging (115→9.4)
Wrapper diag log Ends at 12:19 launch line — no exit code, no diagnostics, no MONITOR_RESULT
Reboot? No (uptime continuous, 2 days 21:45)
Kernel GPU fault (Xid)? None in 12:30–14:05 window (journal readable)
OOM killer (kernel + userspace)? No signature in readable logs
Monitor Detected case-4 (wrapper dead, retries unexhausted), Telegram’d, kept cron
Partial outputs work/hq4-full/ holds only backbone_stub/ — 6 blocks of progress fully lost (AutoRound has no resume)

Lessons (applied to future runs)

  1. pkill -f quantize matches tools/quantize_moss_autoround_awq.py — any peer cleanup for their quant work can kill ours. Use pidfiles + narrow patterns on both sides.
  2. Wrappers should trap SIGTERM and log it — narrows future forensics (SIGKILL stays invisible).
  3. pgrep -f <name> self-matches the monitoring agent’s own cmdline (prompt contains the name) → false “alive”. Use pgrep -f "[n]ame" form plus pidfile checks.
  4. Long AutoRound runs are all-or-nothing — budget for restarts, keep attempts ≥ 2.