Timeline (NZST, machine clock verified via log correlation)
- 12:19 — Full-spec build launched on GPU 0 via
ops/hq4-full/retry-wrapper-hq4-full.sh
(setsid-detached): 256 samples / seqlen 1024 / 200 iters / grad-accum 8 / symmetric
W4A16 G128, target artifacts/MOSS-TTS-v1.5-HQ4-256. Cron monitor
MOSSHQ4FULL_HAIKU_MONITOR installed after a passing manual test.
- 12:19–12:47 — Healthy progress: 6/36 blocks, steady ~280–295 s/block,
per-block loss converging (e.g. 115.2→9.4, 1.66→0.52), peak RAM 45.6 GB,
peak VRAM 19.2 GB. Two monitor checks green, Telegram progress updates sent.
- 12:47:49 — Last log line (
layers.6: 6/36). Process goes silent mid-block-6.
- ~14:01 — Scheduled monitor check finds: wrapper dead, quant pid dead,
GPU 0 idle, no
MONITOR_RESULT, log frozen 73 min. Sends Telegram
needs-attention alert (case 4: wrapper died before exhausting retries),
keeps cron per its instructions. A second attention check follows.
- 14:05 — Human-owned investigation (this entry): rules out reboot (uptime
continuous), kernel GPU faults (Xid-free window 12:30–14:05, journal readable),
kernel + userspace OOM (no signatures). Remaining explanation: external kill
from parallel session activity. Killer unidentified.
- 14:0x — Cron entry removed to stop 15-minute alert spam now that a human
owns the situation (reversible one-liner, recorded here). No relaunch per
explicit user instruction (“don’t quantize nothing else right now”).
Notes
- An earlier
NVRM: Xid 43 at 11:50 is unrelated (predates the 12:19 launch;
hit parity/envelope-era processes, which completed normally).
work/hq4-full/ contains only backbone_stub/ — nothing to salvage or resume.
- Open question for the user/peer: did anyone run
kill/pkill (or a cleanup
matching quantize) between 12:47 and 14:00? Answer would close the case.