Back to the experimentSupporting notebook

MOSS-TTS-v1.5 HQ4 — Phase 1 incident (2026-09-18)

Source: moss-tts-hq4/phase1-incident-2026-09/JOURNEY.md · revision 6550ead3945b

Timeline (NZST, machine clock verified via log correlation)

  1. 12:19 — Full-spec build launched on GPU 0 via ops/hq4-full/retry-wrapper-hq4-full.sh (setsid-detached): 256 samples / seqlen 1024 / 200 iters / grad-accum 8 / symmetric W4A16 G128, target artifacts/MOSS-TTS-v1.5-HQ4-256. Cron monitor MOSSHQ4FULL_HAIKU_MONITOR installed after a passing manual test.
  2. 12:19–12:47 — Healthy progress: 6/36 blocks, steady ~280–295 s/block, per-block loss converging (e.g. 115.2→9.4, 1.66→0.52), peak RAM 45.6 GB, peak VRAM 19.2 GB. Two monitor checks green, Telegram progress updates sent.
  3. 12:47:49 — Last log line (layers.6: 6/36). Process goes silent mid-block-6.
  4. ~14:01 — Scheduled monitor check finds: wrapper dead, quant pid dead, GPU 0 idle, no MONITOR_RESULT, log frozen 73 min. Sends Telegram needs-attention alert (case 4: wrapper died before exhausting retries), keeps cron per its instructions. A second attention check follows.
  5. 14:05 — Human-owned investigation (this entry): rules out reboot (uptime continuous), kernel GPU faults (Xid-free window 12:30–14:05, journal readable), kernel + userspace OOM (no signatures). Remaining explanation: external kill from parallel session activity. Killer unidentified.
  6. 14:0x — Cron entry removed to stop 15-minute alert spam now that a human owns the situation (reversible one-liner, recorded here). No relaunch per explicit user instruction (“don’t quantize nothing else right now”).

Notes