Journal

A coding agent that knows when to ask for help

Three live recovery scenarios passed, each finishing with 10/10 checks.

Vlad / experimentos.
The boundary

The failure triggers were deliberately arranged; this establishes integration behavior, not general coding accuracy.

Measurements

Recorded result

What recovery cost in three live scenarios

Single RTX 3090 · local Qwen worker + GLM reviewer · each scenario finished 10/10 checks

What recovery cost in three live scenariosRepeated failures · 2 reviews: 150 seconds; Acceptance failure · 1 review: 142 seconds; Successful task · no review: 9 seconds. Deliberately triggered integration scenarios; duration does not measure general coding accuracy.Repeated failures · 2 reviews150Repeated failures · 2 reviews: 150 secondsAcceptance failure · 1 review142Acceptance failure · 1 review: 142 secondsSuccessful task · no review9Successful task · no review: 9 seconds0seconds
  1. Repeated failures · 2 reviews150
  2. Acceptance failure · 1 review142
  3. Successful task · no review9

seconds

Deliberately triggered integration scenarios; duration does not measure general coding accuracy.

View data & source
What recovery cost in three live scenarios · seconds
ConfigurationValue
Repeated failures · 2 reviews150
Acceptance failure · 1 review142
Successful task · no review9

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

A compact handoff can still add information

Serialized context at the recovery handoff

A compact handoff can still add informationRepeated failures · before: 14,325 characters; Repeated failures · after: 5,723 characters; Acceptance failure · before: 5,290 characters; Acceptance failure · after: 6,566 characters. The short acceptance request grew after adding diagnostic information; this is not a universal compression ratio.Repeated failures · before14,325Repeated failures · before: 14,325 charactersRepeated failures · after5,723Repeated failures · after: 5,723 charactersAcceptance failure · before5,290Acceptance failure · before: 5,290 charactersAcceptance failure · after6,566Acceptance failure · after: 6,566 characters0characters
  1. Repeated failures · before14,325
  2. Repeated failures · after5,723
  3. Acceptance failure · before5,290
  4. Acceptance failure · after6,566

characters

The short acceptance request grew after adding diagnostic information; this is not a universal compression ratio.

View data & source
A compact handoff can still add information · characters
ConfigurationValue
Repeated failures · before14,325
Repeated failures · after5,723
Acceptance failure · before5,290
Acceptance failure · after6,566

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

A coding agent can keep repeating a failed approach. This experiment gave a local Qwen worker an independent reviewer, then checked whether it could recover without weakening the final acceptance tests.

From the original notebook

Result: all three functional scenarios passed using real Qwen and real GLM.

Model and GPU

  • Worker: Qwen3.8-27B, Unsloth UD-Q4_K_XL, on one RTX 3090: physical GPU 1.
  • Model artifact: Qwen3.8-27B-UD-Q4_K_XL.gguf.
  • Existing model-proxy route: bench-udq4.
  • Existing serving setup retained: 131,072 context, Q8 KV cache, MTP draft length 3. Qwen used medium thinking.
  • Reviewer: actual OpenCode -> Z.AI glm-5.3. All three final-run review requests recorded reasoning_effort=max, thinking.type=enabled, and zero tools.
  • The prior Q4 benchmark had finished. Its GPU was admitted through the existing queue. No occupied GPU job was stopped or changed.
  • The queued backend was assigned to GPU [1] and released automatically after the proxy idle timeout.

End-to-end results

Scenario GLM reviews Final tests Time Result
Repeated failing tool call 2 10/10 passed 150 s Passed
Basic checks pass, independent acceptance fails 1 10/10 passed 142 s Passed
Already-correct implementation 0 10/10 passed 9 s Passed

The two recovery scenarios began with a real Python TTL/LRU cache containing defects: six of ten checks failed. The scenarios deliberately arrange the trigger, so these results establish integration behavior, not general Qwen accuracy. No model replies were mocked.

OMP detected the failures, sent the task/code/errors to GLM, saved a compact handoff, replaced failed history in the next request, and resumed the same Qwen model. Qwen edited the implementation and passed the unchanged tests. Both recovery checkpoints were persisted to the OMP session.

In the repeated-failure rerun, Qwen repeated the reproduction protocol after the first advice despite the handoff saying it was complete. The second automatic review corrected this, and Qwen finished. The two-review limit was respected.

Compaction and the fix found during testing

The first live pass revealed duplicate guidance in the stop-triggered path: both the compacted context and continuation message included the full advice. This was fixed in the installed extension. Continuation now carries a short instruction; the handoff appears once. The reviewer prompt also asks for a concise response.

Final measured request sizes (serialized message characters, not tokens):

Recovery Messages before -> after Characters before -> after
Repeated failures, first review 10 -> 3 14,325 -> 5,723 (60% smaller)
Acceptance-check failure 8 -> 4 5,290 -> 6,566 (24% larger)

The acceptance case replaced old history correctly, but its short original request did not contain the newly discovered regression failure. Adding useful diagnosis and next steps made that next request larger. Compaction therefore does not guarantee a smaller request for every short conversation.

One exploratory check expected every next request to be smaller. The acceptance-check request grew after the failure report and useful advice were added, so that size check did not pass. The functional verification confirmed that OMP removed the earlier tool/assistant history, preserved the user request, and included the guidance once. The published summary records both measurements and this limitation.

Validation and installation

  • Nine focused recovery unit/process tests passed.
  • Six compiled OMP CLI integration scenarios passed, including a regression assertion that guidance appears once.
  • TypeScript checking passed.
  • Installed OMP binary and extension checksums were verified.
  • The core OMP binary did not change during this test. The extension was updated; the previous extension and manifest are backed up under backups/extension-20261001T023732Z/.
  • The original pre-patch OMP backup and restore manifest were preserved. Open a new OMP session to load the revised extension.

A broad bun test test filter initially traversed the bundled upstream checkout and hit an unrelated watchdog test failure. That scan was stopped. The test command was corrected to explicit local test files; the nine intended tests passed. The temporary watchdog fixture was removed.

Evidence

The verification summary records the model configuration, scenario outcomes, review settings, and measured context sizes. Full transcripts and prompts remain in the author workstation run archive and are not copied into this repository. The published summary excludes credentials and workstation paths.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS