Back to the experimentSupporting notebook

OMP + GLM recovery: live Qwen UD Q4 experiment

Source: agentic-workflows/omp-glm-recovery-qwen-udq4-2026-10-01/README.md · revision 6550ead3945b

Result: all three functional scenarios passed using real Qwen and real GLM.

Model and GPU

End-to-end results

Scenario GLM reviews Final tests Time Result
Repeated failing tool call 2 10/10 passed 150 s Passed
Basic checks pass, independent acceptance fails 1 10/10 passed 142 s Passed
Already-correct implementation 0 10/10 passed 9 s Passed

The two recovery scenarios began with a real Python TTL/LRU cache containing defects: six of ten checks failed. The scenarios deliberately arrange the trigger, so these results establish integration behavior, not general Qwen accuracy. No model replies were mocked.

OMP detected the failures, sent the task/code/errors to GLM, saved a compact handoff, replaced failed history in the next request, and resumed the same Qwen model. Qwen edited the implementation and passed the unchanged tests. Both recovery checkpoints were persisted to the OMP session.

In the repeated-failure rerun, Qwen repeated the reproduction protocol after the first advice despite the handoff saying it was complete. The second automatic review corrected this, and Qwen finished. The two-review limit was respected.

Compaction and the fix found during testing

The first live pass revealed duplicate guidance in the stop-triggered path: both the compacted context and continuation message included the full advice. This was fixed in the installed extension. Continuation now carries a short instruction; the handoff appears once. The reviewer prompt also asks for a concise response.

Final measured request sizes (serialized message characters, not tokens):

Recovery Messages before -> after Characters before -> after
Repeated failures, first review 10 -> 3 14,325 -> 5,723 (60% smaller)
Acceptance-check failure 8 -> 4 5,290 -> 6,566 (24% larger)

The acceptance case replaced old history correctly, but its short original request did not contain the newly discovered regression failure. Adding useful diagnosis and next steps made that next request larger. Compaction therefore does not guarantee a smaller request for every short conversation.

One exploratory check expected every next request to be smaller. The acceptance-check request grew after the failure report and useful advice were added, so that size check did not pass. The functional verification confirmed that OMP removed the earlier tool/assistant history, preserved the user request, and included the guidance once. The published summary records both measurements and this limitation.

Validation and installation

A broad bun test test filter initially traversed the bundled upstream checkout and hit an unrelated watchdog test failure. That scan was stopped. The test command was corrected to explicit local test files; the nine intended tests passed. The temporary watchdog fixture was removed.

Evidence

The verification summary records the model configuration, scenario outcomes, review settings, and measured context sizes. Full transcripts and prompts remain in the author workstation run archive and are not copied into this repository. The published summary excludes credentials and workstation paths.