Skip to content

Scale the txn_box multi-ramp autest tolerance to the binomial spread - #13726

Merged
bryancall merged 1 commit into
apache:masterfrom
bryancall:txn_box-ramp-binomial-tolerance
Sep 28, 2026
Merged

bryancall merged 1 commit into
apache:masterfrom
bryancall:txn_box-ramp-binomial-tolerance

Conversation

@bryancall

@bryancall bryancall commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary: a test that flaked about once a week should now go more than a decade

Before After
Chance a working build fails a CI run, by bad luck alone about 1 in 340 about 1 in 210,000
At about 40 CI runs a day about once a week about once every 14 years

Two bell curves, one per ramp bucket, showing where a working build's count lands. The old pass range is a fixed 50 requests either side: only 3.45 standard deviations for the 30% bucket, 5.3 for the 10% bucket. The new range is 5 standard deviations either side for both.

The txn_box multi-ramp autests send 1000 requests through a random percentage ramp and count how many land in each bucket, and that count wobbles from run to run. The old pass range was a flat plus or minus 5 percentage points: roomy for the 10% bucket, tight for the 30% ones. This sizes each range to how much its bucket actually wobbles. A broken ramp still fails, and only the test's checker changes: no plugin or server code.

The CI rate is measured from recent Jenkins autest builds over a few busy hours, so the real rate is probably lower and the gap between flakes longer.

Fixed: the 30% buckets no longer fail a working build by chance

Bucket Old range New range Chance a working build fails, old New
30% 250..350 227..373 about 1 in 2,000 about 1 in 2.3 million
10% 50..150 52..148 about 1 in 3.6 million about 1 in 1.4 million
A whole CI run: txn_box_multi-ramp-1, -2 and -3, six 30% buckets and three 10% about 1 in 340 about 1 in 210,000

Each transaction draws its own random number, so a bucket count is binomial, with a standard deviation of sqrt(RepeatCount * p * (1 - p)): 14.5 at 30% and 9.5 at 10%, for 1000 requests. A flat plus or minus 5 point window therefore sits 3.45 standard deviations out at 30% but 5.3 at 10%, which is why only the 30% buckets ever flaked.

This change sizes the window as 5 standard deviations either side of the expected count, so every bucket gets the same margin. Targets 0 and 100 have a standard deviation of zero, so their window stays exact. A failure now also reports the expected count and how many standard deviations away the observed count was.

Verified: a bake-off in the CI image

Looping all three multi-ramp tests in the CI image (ci.trafficserver.apache.org/ats/fedora:42, GCC) on two 32-core hosts. The first round ran the old checker on one host and the new one on the other. The second round ran both side by side on each host, against the same traffic_server binary. Each run's bucket counts were recorded, so both checkers can also be scored against identical traffic.

Test runs Failures
Old checker, run for real 3,273 1 ('v1/video/alias' failed with 249 not in 250..350)
New checker, run for real 3,279 0
Old checker, scored against every recorded run 6,648 5 (0.075%, predicted 0.098%)
New checker, scored against every recorded run 6,648 0 (predicted 0.00016%)

Every failure was a 30% bucket landing just outside 250..350, at counts of 248, 249, 352, 353 and 355. Every run logged all 6,000 transactions, so none of the failures came from a timeout or a truncated log. The observed standard deviations were 14.3 and 14.6 for the two 30% buckets and 9.45 for the 10% bucket, against the binomial prediction of 14.5 and 9.5.

Verified: a broken ramp still fails

Changing multi-ramp-2.cfg.yaml to ramp the 30% buckets at 20% makes the new checker fail them, at 6.3 and 7.0 standard deviations below the expected count:

'v1/video/search' failed with 208 not in 227..373 (expected 300, -6.3 sigma)
'v1/video/alias' failed with 198 not in 227..373 (expected 300, -7.0 sigma)

Not changed: clang builds fail these tests on every run

That is an unrelated bug, the lt comparison rejecting integers, and it is fixed in #13724.

Each transaction draws its own random number, so a bucket count is
binomial and a flat plus or minus 5 point window is a different number
of standard deviations at every target: 5.3 sigma at target 10 but only
3.45 sigma at target 30, which fails about once in 2000 runs per bucket
and once in 340 across the three tests. Size the window from
sqrt(n * p * (1 - p)) instead, so every bucket gets the same margin and
the window tightens as RepeatCount grows.
Copilot AI lite review requested due to automatic review settings September 23, 2026 23:02
@bryancall bryancall added this to the 11.0.0 milestone Sep 23, 2026
@bryancall bryancall added Tests TxnBox TxnBox plugin labels Sep 23, 2026
@bryancall bryancall self-assigned this Sep 23, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The reviewed change addresses flaky validation while preserving deterministic bounds and adding useful diagnostics.

Review effort: Lite
Findings: None

What changed in this PR

Updates txn_box multi-ramp AuTests to use binomial 5-sigma tolerances instead of fixed percentage windows.

Changes:

  • Computes statistically scaled bounds per ramp target.
  • Preserves exact validation for 0% and 100% targets.
  • Reports expected counts and sigma distance on failures.
File Description
tests/​gold_tests/​pluginTest/​txn_box/​ramp/​multi_ramp_common.py Implements statistical tolerance calculations and improved failure diagnostics.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@bryancall
bryancall requested a review from bneradt September 24, 2026 02:51

@bneradt bneradt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed d2b85dc. No actionable findings. Checked the shared validator against the three multi-ramp configurations and the random extractor. Ran 16 focused Python checks against the patched validator covering inclusive boundaries, out-of-range failures, exact 0%/100% expectations, and diagnostics. Independently calculated the binomial tail probabilities, which agree with the stated per-bucket estimates. The wider 30% window trades some sensitivity to small routing biases for fewer random failures. I did not run the full AuTests.

@bneradt bneradt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved.

If I understand this correctly, this changes the tolerance from +- 5% from the ramp setting to verifying that the ramp behavior is within 5 standard deviations of the target. When you commit, can you please ensure that is clear? The current commit message gives some formula which isn't obvious to me that is what it is doing.

@bryancall
bryancall merged commit 9a08e84 into apache:master Sep 28, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Tests TxnBox TxnBox plugin

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants