The question
A coding agent can keep repeating a failed approach. This experiment gave a local Qwen worker an independent reviewer, then checked whether it could recover without weakening the final acceptance tests.
From the original notebook
Result: all three functional scenarios passed using real Qwen and real GLM.
Model and GPU
- Worker: Qwen3.8-27B, Unsloth UD-Q4_K_XL, on one RTX 3090: physical GPU 1.
- Model artifact:
Qwen3.8-27B-UD-Q4_K_XL.gguf. - Existing model-proxy route:
bench-udq4. - Existing serving setup retained: 131,072 context, Q8 KV cache, MTP draft length 3. Qwen used medium thinking.
- Reviewer: actual OpenCode -> Z.AI
glm-5.3. All three final-run review requests recordedreasoning_effort=max,thinking.type=enabled, and zero tools. - The prior Q4 benchmark had finished. Its GPU was admitted through the existing queue. No occupied GPU job was stopped or changed.
- The queued backend was assigned to GPU
[1]and released automatically after the proxy idle timeout.
End-to-end results
| Scenario | GLM reviews | Final tests | Time | Result |
|---|---|---|---|---|
| Repeated failing tool call | 2 | 10/10 passed | 150 s | Passed |
| Basic checks pass, independent acceptance fails | 1 | 10/10 passed | 142 s | Passed |
| Already-correct implementation | 0 | 10/10 passed | 9 s | Passed |
The two recovery scenarios began with a real Python TTL/LRU cache containing defects: six of ten checks failed. The scenarios deliberately arrange the trigger, so these results establish integration behavior, not general Qwen accuracy. No model replies were mocked.
OMP detected the failures, sent the task/code/errors to GLM, saved a compact handoff, replaced failed history in the next request, and resumed the same Qwen model. Qwen edited the implementation and passed the unchanged tests. Both recovery checkpoints were persisted to the OMP session.
In the repeated-failure rerun, Qwen repeated the reproduction protocol after the first advice despite the handoff saying it was complete. The second automatic review corrected this, and Qwen finished. The two-review limit was respected.
Compaction and the fix found during testing
The first live pass revealed duplicate guidance in the stop-triggered path: both the compacted context and continuation message included the full advice. This was fixed in the installed extension. Continuation now carries a short instruction; the handoff appears once. The reviewer prompt also asks for a concise response.
Final measured request sizes (serialized message characters, not tokens):
| Recovery | Messages before -> after | Characters before -> after |
|---|---|---|
| Repeated failures, first review | 10 -> 3 | 14,325 -> 5,723 (60% smaller) |
| Acceptance-check failure | 8 -> 4 | 5,290 -> 6,566 (24% larger) |
The acceptance case replaced old history correctly, but its short original request did not contain the newly discovered regression failure. Adding useful diagnosis and next steps made that next request larger. Compaction therefore does not guarantee a smaller request for every short conversation.
One exploratory check expected every next request to be smaller. The acceptance-check request grew after the failure report and useful advice were added, so that size check did not pass. The functional verification confirmed that OMP removed the earlier tool/assistant history, preserved the user request, and included the guidance once. The published summary records both measurements and this limitation.
Validation and installation
- Nine focused recovery unit/process tests passed.
- Six compiled OMP CLI integration scenarios passed, including a regression assertion that guidance appears once.
- TypeScript checking passed.
- Installed OMP binary and extension checksums were verified.
- The core OMP binary did not change during this test. The extension was updated; the previous extension and manifest are backed up under
backups/extension-20261001T023732Z/. - The original pre-patch OMP backup and restore manifest were preserved. Open a new OMP session to load the revised extension.
A broad bun test test filter initially traversed the bundled upstream checkout and hit an unrelated watchdog test failure. That scan was stopped. The test command was corrected to explicit local test files; the nine intended tests passed. The temporary watchdog fixture was removed.
Evidence
The verification summary records the model configuration, scenario outcomes, review settings, and measured context sizes. Full transcripts and prompts remain in the author workstation run archive and are not copied into this repository. The published summary excludes credentials and workstation paths.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.