Scale the txn_box multi-ramp autest tolerance to the binomial spread - #13726
Conversation
Each transaction draws its own random number, so a bucket count is binomial and a flat plus or minus 5 point window is a different number of standard deviations at every target: 5.3 sigma at target 10 but only 3.45 sigma at target 30, which fails about once in 2000 runs per bucket and once in 340 across the three tests. Size the window from sqrt(n * p * (1 - p)) instead, so every bucket gets the same margin and the window tightens as RepeatCount grows.
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The reviewed change addresses flaky validation while preserving deterministic bounds and adding useful diagnostics.
Review effort: Lite
Findings: None
What changed in this PR
Updates txn_box multi-ramp AuTests to use binomial 5-sigma tolerances instead of fixed percentage windows.
Changes:
- Computes statistically scaled bounds per ramp target.
- Preserves exact validation for 0% and 100% targets.
- Reports expected counts and sigma distance on failures.
| File | Description |
|---|---|
tests/gold_tests/pluginTest/txn_box/ramp/multi_ramp_common.py |
Implements statistical tolerance calculations and improved failure diagnostics. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
bneradt
left a comment
There was a problem hiding this comment.
Reviewed d2b85dc. No actionable findings. Checked the shared validator against the three multi-ramp configurations and the random extractor. Ran 16 focused Python checks against the patched validator covering inclusive boundaries, out-of-range failures, exact 0%/100% expectations, and diagnostics. Independently calculated the binomial tail probabilities, which agree with the stated per-bucket estimates. The wider 30% window trades some sensitivity to small routing biases for fewer random failures. I did not run the full AuTests.
bneradt
left a comment
There was a problem hiding this comment.
Approved.
If I understand this correctly, this changes the tolerance from +- 5% from the ramp setting to verifying that the ramp behavior is within 5 standard deviations of the target. When you commit, can you please ensure that is clear? The current commit message gives some formula which isn't obvious to me that is what it is doing.
Summary: a test that flaked about once a week should now go more than a decade
The txn_box multi-ramp autests send 1000 requests through a random percentage ramp and count how many land in each bucket, and that count wobbles from run to run. The old pass range was a flat plus or minus 5 percentage points: roomy for the 10% bucket, tight for the 30% ones. This sizes each range to how much its bucket actually wobbles. A broken ramp still fails, and only the test's checker changes: no plugin or server code.
The CI rate is measured from recent Jenkins autest builds over a few busy hours, so the real rate is probably lower and the gap between flakes longer.
Fixed: the 30% buckets no longer fail a working build by chance
txn_box_multi-ramp-1,-2and-3, six 30% buckets and three 10%Each transaction draws its own random number, so a bucket count is binomial, with a standard deviation of
sqrt(RepeatCount * p * (1 - p)): 14.5 at 30% and 9.5 at 10%, for 1000 requests. A flat plus or minus 5 point window therefore sits 3.45 standard deviations out at 30% but 5.3 at 10%, which is why only the 30% buckets ever flaked.This change sizes the window as 5 standard deviations either side of the expected count, so every bucket gets the same margin. Targets 0 and 100 have a standard deviation of zero, so their window stays exact. A failure now also reports the expected count and how many standard deviations away the observed count was.
Verified: a bake-off in the CI image
Looping all three multi-ramp tests in the CI image (
ci.trafficserver.apache.org/ats/fedora:42, GCC) on two 32-core hosts. The first round ran the old checker on one host and the new one on the other. The second round ran both side by side on each host, against the sametraffic_serverbinary. Each run's bucket counts were recorded, so both checkers can also be scored against identical traffic.'v1/video/alias' failed with 249 not in 250..350)Every failure was a 30% bucket landing just outside 250..350, at counts of 248, 249, 352, 353 and 355. Every run logged all 6,000 transactions, so none of the failures came from a timeout or a truncated log. The observed standard deviations were 14.3 and 14.6 for the two 30% buckets and 9.45 for the 10% bucket, against the binomial prediction of 14.5 and 9.5.
Verified: a broken ramp still fails
Changing
multi-ramp-2.cfg.yamlto ramp the 30% buckets at 20% makes the new checker fail them, at 6.3 and 7.0 standard deviations below the expected count:Not changed: clang builds fail these tests on every run
That is an unrelated bug, the
ltcomparison rejecting integers, and it is fixed in #13724.