Adding authentication, path confinement, and worker-identity checks to a GPU model proxy — then letting the test suite pick the production version. The hardened build won 107 to 100 with zero regressions and a 7-for-7 control-arm confirmation.
Executive summary
- The hardened proxy tip passes 107 tests, 1 skipped, in 0.37 s; the pre-hardening tip passes 100 tests, 1 skipped, in 0.38 s.
- All 7 new hardening tests fail against pre-hardening code and pass against hardened code, confirming they measure the change rather than the harness.
- All 100 pre-existing tests still pass on the hardened build: zero regressions, suite time unchanged within noise (−2.6%).
- Recommended: serve the hardened
maintip (168d60d); it is behaviorally identical to the validated live tip plus one docs file. - Principal caveat: validation is unit-level only — no live backend, GPU load, or real media job was run.
The problem (why)
A local inference proxy routes chat, image, upscale, and speech requests across a farm of GPU backends. That convenience concentrates risk: one unauthenticated route or one unvalidated file path can turn the proxy into an open relay or a disk-writing primitive. The branch under test layered a hardening pass onto the live proxy — bearer auth for the upscale endpoint, filesystem confinement for media staging, host allowlisting for remote inputs, identity checks before adopting stray workers, and telemetry for previously unmetered routes.
The operational question was binary: promote the hardened branch to main, or hold it back? Three candidate states existed — pre-hardening live tip, hardened live tip, and the main merge result — and the decision needed recorded evidence, not a reading of the diff.
The system and experiment (what)
The system is a FastAPI proxy (Python 3.12) fronting llama.cpp, vLLM, and ComfyUI backends on a 3× RTX 3090 workstation. The test suite is the measurement instrument: 19 files, 107 test functions (one parametrized ×2, hence 108 collected cases), run with AUTO_DISABLE_MISSING_WEIGHTS=0 so verdicts do not depend on which weights happen to be on disk.
The variable under test is the code state: control A (pre-hardening live tip ccf2f53), treatment B (hardened live tip 27f3620, i.e. A plus the 5-file hardening layer and a 5-line auth fix), and ship candidate C (B merged into main as 168d60d). A fourth run — the 7 new tests executed against pre-hardening code — serves as the control arm proving the new tests discriminate.
Method (how)
Each arm ran the same command from the proxy checkout: pytest tests/ -v --durations=25, plus a JUnit export and a --collect-only inventory. Timing is pytest’s reported execution time (sum of test phases, not wall clock); wall time was separately observed at ~0.58 s for the full run.
Validation came in layers, not just one green run. The pre-hardening tip was tested before any merge (baseline). The hardening diff was reviewed item by item against the code before merging. The merge itself was verified conflict-free. The post-merge tip was tested, then the main merge result was tested again in an isolated checkout. Finally, the control arm ran the new tests against the old code to confirm they fail without the treatment.
Jargon, briefly: a parametrized test is one test function the runner executes multiple times with different inputs (hence 107 functions → 108 cases). A control arm is the untreated variant run to prove the measurement responds to the treatment. Guardrail metrics are quantities that must not get worse — here, the pre-existing pass count and suite time.
Results
The decision-relevant table first — absolute verdicts per arm, no sampling involved:
| Arm | Suite verdict | New tests | Pre-existing passing | pytest time |
|---|---|---|---|---|
| A pre-hardening | 100 passed, 1 skipped | n/a (absent) | 100 | 0.38 s |
| B hardened | 107 passed, 1 skipped | 7 passed | 100 | 0.37 s |
| C main merge | 107 passed, 1 skipped | 7 passed | 100 | 0.35–0.37 s |
| Control (new tests on A) | — | 7 failed | — | 0.24 s |
Coverage gain: 107 − 100 = 7 tests (+7.0%). Timing delta: (0.37 − 0.38) / 0.38 = −2.6%, inside run-to-run noise and well within the ±20% guardrail. The single skip in every arm is environmental — a weights-absent conditional skip for one elastic-route test — and identical across arms, so it cannot favor any version.
Figure 1 data (source: data/qualification.json, runs[].result):
version,passed,skipped
A pre-hardening (ccf2f53),100,1
B hardened (27f3620),107,1
C main merge (168d60d),107,1
Figure 1. Chart specification: horizontal bar chart, one bar per version, x-axis “tests passed” (unit: count, range 0–110), bars labeled with exact values, sorted A → B → C. Caption: hardened builds pass 7 more tests than pre-hardening with identical skips. Alt text: three horizontal bars showing 100, 107, and 107 tests passed for versions A, B, and C. Readers should notice the +7 step appears exactly once, at the hardening layer — the main merge adds no tests. The chart cannot establish runtime safety or performance; it counts unit verdicts only.
Why the result looks this way
The +7 step is directly observed: the hardening layer added one test file with 7 functions covering auth rejection (503 unconfigured, 401 wrong token), path-escape blocking, default-deny remote hosts, adoption refusal for unverified listeners, opt-in router force-switches, and VibeVoice request accounting (see data/pr5-tests-verbose.log). The control arm shows all 7 failing on old code with missing-attribute errors — the mechanism is visible, not inferred.
The zero-regression half is equally direct: the 100 pre-existing verdicts are unchanged across arms (same names in data/pytest-full-verbose.log and data/test-inventory.txt). A plausible interpretation — offered as hypothesis, not measurement — is that the hardening touched mostly additive seams (new auth helpers, new guards, new metrics calls) rather than rewriting dispatch logic, which is why nothing old broke. The flat suite time follows from the tests being pure unit checks with no network or GPU waits; slowest observed case was 0.03 s.
Operational recommendation
Serve version C, the main tip. It is the only arm that is both hardened and promoted, and its delta versus validated B is a single docs file, so no behavioral divergence is possible between what was tested and what ships.
Trade-offs and rollout: the hardening contains three intended breaking changes — the upscale route now requires a bearer token (HTTP 503 until configured), remote media inputs require an explicit host allowlist (default deny), and relative upscale outputs now resolve under the ComfyUI output root. Before routing production traffic: set the token and allowlist from the process environment, then smoke-test with one authenticated upscale call and one small chat completion, and confirm the new metrics endpoints record both requests.
Limits and next experiment
Comparison boundaries, stated plainly. The suite never starts a backend, touches a GPU, or processes real media — it validates config, routing, guards, and accounting logic. GPU/VRAM behavior, real upscale quality, VibeVoice container cycling, and end-to-end latency were reviewed in code only. The verdict “hardened wins” is therefore scoped to specified behavior under unit test, not to production robustness.
What this does not prove: that the proxy is secure (no adversarial testing was performed), that it is faster (no benchmark ran), or that the hardening is complete (it covers the named surfaces only).
The smallest controlled follow-up: stage one real upscale job and one speech request against the hardened tip and confirm the new metrics rows, the auth gates, and the idle shutdown fire end to end. That single experiment would convert the strongest remaining claim — “the guards work in operation” — from reviewed to measured.
How to reproduce or audit
Portable checks (any clone of the proxy repo with the pinned toolchain):
./scripts/rerun-suite.sh /path/to/proxy-checkout
This regenerates the verbose log, durations, JUnit report, inventory, and skip summary from the experiment’s scripts/ directory. Toolchain: Python 3.12.13, pytest 9.1.1 (see data/versions.txt).
Source-workstation-only caveat: one route test reads a host-local launcher script that is intentionally untracked in git; a bare clone fails exactly that test with FileNotFoundError until the host file is staged. This was observed once during the experiment and is recorded here so auditors do not mistake it for a regression.
Evidence appendix
Source files (all under the experiment directory): README.md (full validation report), CONCLUSION.md (version comparison), SESSION.md (session narrative), data/qualification.json (machine-readable runs), data/pytest-full-verbose.log (108 verdicts), data/pr5-tests-verbose.log (7 new tests), data/control-arm.log (7 failures on old code), data/pytest-junit.xml, data/test-inventory.txt, data/git-evidence.txt, data/versions.txt, data/skip-reason.txt, data/main-merge-sha.txt, scripts/rerun-suite.sh.
flowchart TD
A[Baseline suite on A<br/>100 passed] --> B[Review hardening diff<br/>13 items checked]
B --> C[Merge review branch<br/>zero conflicts]
C --> D[Fix review finding<br/>API-key forwarding]
D --> E[Suite on B<br/>107 passed]
E --> F[Suite on C (main)<br/>107 passed]
E --> G[Control arm: new tests on A<br/>7 failed]
F --> H{PROMOTE?}
G --> H
H -->|all gates green| I[Ship C]
Figure 2. Validation funnel that gated the promotion decision; every box is a recorded run or review in this experiment’s data. Alt text: flowchart from baseline suite through review, merge, fix, two green suites, and a failing control arm, converging on the decision to ship version C. Readers should notice the control arm is what makes the +7 meaningful — without it, new passing tests could be vacuous. The diagram cannot establish what untested behavior (GPU, media, containers) does in production.
Glossary: live tip — head of the staging branch under test; hardening layer — the 5-file auth/confinement/identity/metrics change set; control arm — new tests executed against pre-hardening code; guardrail — a metric that must not regress; parametrized — one test function run with multiple inputs.