Back to the experimentSupporting notebook

Conclusion: which proxy version wins — pre-hardening vs hardened vs main

Source: qwen38/proxy-pr5-hardening-validation-2026-09-14/CONCLUSION.md · revision 6550ead3945b

Decision question: of the three validated proxy states, which should serve production traffic? Comparison design follows a fixed hypothesis with a primary metric and guardrails; all numbers below are measured (see data/).

Hypothesis

Because the unhardened proxy exposes an unauthenticated upscale endpoint, unconfined file staging, unverified worker adoption, and unmetered auxiliary routes, we believe the PR #5 hardened version is strictly safer to operate than the pre-hardening version, with zero regressions and no suite slowdown. We will know this is true when: (primary) the hardened tip passes the full suite including 7 new covering tests; (guardrails) all 100 pre-existing tests still pass and suite time stays within ±20% of baseline.

Variants

Arm Commit Description
A (control) ccf2f53 Live tip before PR #5 (proxy-managed :12702 only)
B (treatment) 27f3620 A + PR #5 hardening + LLAMA_API_KEY identity fix
C (ship candidate) 168d60d B merged into main (PR #3 history retained)
Control-arm check ccf2f53 + new test file New tests run against pre-hardening code

A and B differ only by the hardening layer (+552/-43 across 5 files, plus a 5-line auth fix). C differs from B by one docs file only (docs/request-metrics.md), so C is expected to behave identically to B.

Results

Metric A (ccf2f53) B (27f3620) C (168d60d)
Full suite 100 passed, 1 skipped 107 passed, 1 skipped 107 passed, 1 skipped
New hardening tests n/a (file absent) 7 passed 7 passed
Pre-existing tests passing 100 100 (zero regressions) 100 (zero regressions)
Suite time (pytest) 0.38 s 0.37 s 0.35–0.37 s
Control-arm check: new tests on A — 7 failed (all) —

Coverage gain formula:

new_coverage = tests(B) - tests(A) = 107 - 100 = 7 tests (+7.0%)

Timing delta: (0.37 - 0.38) / 0.38 = -2.6% — inside noise, well within the ±20% guardrail.

The control-arm check is the decisive validity evidence: the same 7 test functions fail 7/7 against pre-hardening code (AttributeError on the missing hardening symbols) and pass 7/7 against hardened code. The tests measure the treatment, not the harness (control-arm.log).

Verdict: ship the hardened version (C, identical to B)

Winner: 168d60d on main. It carries the full hardening layer, all guardrails are green, and it is byte-identical in behavior to the validated live tip. Runner-up B is equally good technically; C wins only because it is the promoted, retained-history ref. Control A loses: it serves an unauthenticated upscale route and unconfined staging with no covering tests.

Per-area secondary confirmation (all measured on B/C):

What could change this verdict

Limits

Deterministic unit suite, not sampled traffic: there is no p-value here, and none is needed — outcomes are pass/fail per commit, reproduced across runs. What this does not prove: runtime performance, GPU behavior, real upscale quality, or VibeVoice container cycling. Smallest controlled follow-up: stage one real upscale job and one VibeVoice ASR request against B and confirm the new metrics rows and auth gates fire end to end.