Decision question: of the three validated proxy states, which should serve production traffic? Comparison design follows a fixed hypothesis with a primary metric and guardrails; all numbers below are measured (see data/).
Hypothesis
Because the unhardened proxy exposes an unauthenticated upscale endpoint, unconfined file staging, unverified worker adoption, and unmetered auxiliary routes, we believe the PR #5 hardened version is strictly safer to operate than the pre-hardening version, with zero regressions and no suite slowdown. We will know this is true when: (primary) the hardened tip passes the full suite including 7 new covering tests; (guardrails) all 100 pre-existing tests still pass and suite time stays within ±20% of baseline.
Variants
| Arm | Commit | Description |
|---|---|---|
| A (control) | ccf2f53 |
Live tip before PR #5 (proxy-managed :12702 only) |
| B (treatment) | 27f3620 |
A + PR #5 hardening + LLAMA_API_KEY identity fix |
| C (ship candidate) | 168d60d |
B merged into main (PR #3 history retained) |
| Control-arm check | ccf2f53 + new test file |
New tests run against pre-hardening code |
A and B differ only by the hardening layer (+552/-43 across 5 files, plus a
5-line auth fix). C differs from B by one docs file only
(docs/request-metrics.md), so C is expected to behave identically to B.
Results
| Metric | A (ccf2f53) |
B (27f3620) |
C (168d60d) |
|---|---|---|---|
| Full suite | 100 passed, 1 skipped | 107 passed, 1 skipped | 107 passed, 1 skipped |
| New hardening tests | n/a (file absent) | 7 passed | 7 passed |
| Pre-existing tests passing | 100 | 100 (zero regressions) | 100 (zero regressions) |
| Suite time (pytest) | 0.38 s | 0.37 s | 0.35–0.37 s |
| Control-arm check: new tests on A | — | 7 failed (all) | — |
Coverage gain formula:
new_coverage = tests(B) - tests(A) = 107 - 100 = 7 tests (+7.0%)
Timing delta: (0.37 - 0.38) / 0.38 = -2.6% — inside noise, well within the
±20% guardrail.
The control-arm check is the decisive validity evidence: the same 7 test
functions fail 7/7 against pre-hardening code (AttributeError on the
missing hardening symbols) and pass 7/7 against hardened code. The tests
measure the treatment, not the harness (control-arm.log).
Verdict: ship the hardened version (C, identical to B)
Winner: 168d60d on main. It carries the full hardening layer, all
guardrails are green, and it is byte-identical in behavior to the validated
live tip. Runner-up B is equally good technically; C wins only because it is
the promoted, retained-history ref. Control A loses: it serves an
unauthenticated upscale route and unconfined staging with no covering tests.
Per-area secondary confirmation (all measured on B/C):
- Upscale auth rejects unconfigured (503) and wrong-token (401) callers.
- Path roots block directory escape; remote inputs default-deny without an explicit host allowlist.
- Elastic pool refuses to adopt unverified listeners.
- Router sweeps no longer force-switch unless
--forceis passed. - VibeVoice request accounting balances (+1/−1 with completion timestamp).
What could change this verdict
- A live-fire test showing the hardening breaks a real workload (none run; validation is unit-level only).
- Discovery that an operator depends on unauthenticated
/v1/upscaleor unallowlisted remote inputs — the intended breaking changes (see README operational notes). If such a dependency exists, the verdict stands but the rollout needs a token/allowlist migration first.
Limits
Deterministic unit suite, not sampled traffic: there is no p-value here, and none is needed — outcomes are pass/fail per commit, reproduced across runs. What this does not prove: runtime performance, GPU behavior, real upscale quality, or VibeVoice container cycling. Smallest controlled follow-up: stage one real upscale job and one VibeVoice ASR request against B and confirm the new metrics rows and auth gates fire end to end.