Skip to content

Implement dynamic auto test retries - #12467

Merged
gh-worker-dd-mergequeue-cf854d[bot] merged 12 commits into
masterfrom
gnufede/dynamic-atr-retries
Sep 23, 2026
Merged

gh-worker-dd-mergequeue-cf854d[bot] merged 12 commits into
masterfrom
gnufede/dynamic-atr-retries

Conversation

@gnufede

@gnufede gnufede commented Sep 11, 2026

Copy link
Copy Markdown
Member

Motivation

Add dynamic Auto Test Retries (ATR) budgets based on test duration, instead of the flat per-test retry limit. When enabled, the number of retries allowed for a test is determined by the duration of its initial attempt, using the same duration buckets as Early Flake Detection (5s / 10s / 30s / 5m / >5m).

Two new env vars control the behavior:

  • DD_CIVISIBILITY_DYNAMIC_ATR_ENABLED — enables duration-based ATR budgets. When unset/false, ATR stays on its existing flat-limit path.
  • DD_CIVISIBILITY_DYNAMIC_ATR_BUCKETS — optionally overrides the five duration-based retry budgets with five positive comma-separated integers in [1, 20]. When unset/empty, the EFD retry settings from the backend are used.

The new handler classifies each test once (by its initial-attempt duration) and caches the resulting max-retries count for the lifetime of that test. The duration bucket index is computed via shared helpers on the EFD settings type, so bucket boundaries stay consistent with EFD (the existing EFD retry handler is refactored to use them too).

Telemetry: records a dynamic_atr_retries.enabled count metric with a has_custom_buckets tag when dynamic ATR is enabled.

Feature parity with dd-trace-py PR #20028.

Add dynamic Auto Test Retries (ATR) budgets based on test duration,
instead of the flat per-test retry limit. When enabled, the number of
retries allowed for a test is determined by the duration of its initial
attempt, using the same duration buckets as Early Flake Detection
(5s / 10s / 30s / 5m / >5m).

Two new env vars control the behavior:
- DD_CIVISIBILITY_DYNAMIC_ATR_ENABLED — enables duration-based ATR
  budgets. When unset/false, ATR stays on its existing flat-limit path.
- DD_CIVISIBILITY_DYNAMIC_ATR_BUCKETS — optionally overrides the five
  duration-based retry budgets with five positive comma-separated
  integers in [1, 20]. When unset/empty, the EFD retry settings from
  the backend are used.

The new handler classifies each test once (by its initial-attempt
duration) and caches the resulting max-retries count for the lifetime
of that test. The duration bucket index is computed via shared helpers
on the EFD settings type, so bucket boundaries stay consistent with EFD
(the existing EFD retry handler is refactored to use them too).

Telemetry: records a `dynamic_atr_retries.enabled` count metric with a
`has_custom_buckets` tag when dynamic ATR is enabled.

Feature parity with dd-trace-py PR #20028.
@datadog-official

This comment has been minimized.

@dd-octo-sts

dd-octo-sts Bot commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 14.10 s 13.94 s [+0.4%; +1.9%] (maybe worse)
startup:insecure-bank:tracing:Agent 12.93 s 12.98 s [-1.3%; +0.5%] (no difference)
startup:petclinic:appsec:Agent 17.09 s 16.86 s [+0.4%; +2.4%] (maybe worse)
startup:petclinic:iast:Agent 16.94 s 17.03 s [-1.5%; +0.5%] (no difference)
startup:petclinic:profiling:Agent 16.51 s 16.40 s [-3.9%; +5.3%] (no difference)
startup:petclinic:sca:Agent 16.97 s 16.66 s [+0.7%; +3.1%] (maybe worse)
startup:petclinic:tracing:Agent 16.12 s 16.08 s [-0.5%; +1.1%] (no difference)

Commit: 63303ba8 · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

@daniel-mohedano daniel-mohedano added type: feature Enhancements and improvements comp: ci visibility Continuous Integration Visibility tag: ai generated Largely based on code generated by an AI or LLM labels Sep 18, 2026
@daniel-mohedano daniel-mohedano changed the title feat(ci_visibility): dynamic atr retries Implement dynamic auto test retries Sep 18, 2026
@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - nebula-release-plugin

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 36.45 36.42 $\color{red}{\blacktriangle}$ +0.03 36.42 $\color{red}{\blacktriangle}$ +0.03 34/113
agentless 35.36 34.99 $\color{red}{\blacktriangle}$ +0.37 35.70 $\color{green}{\blacktriangledown}$ -0.34 34/113
agentlessCodeCoverage 44.61 44.48 $\color{red}{\blacktriangle}$ +0.13 44.48 $\color{red}{\blacktriangle}$ +0.13 34/113
agentlessLineCoverage 55.89 55.43 $\color{red}{\blacktriangle}$ +0.46 55.43 $\color{red}{\blacktriangle}$ +0.46 35/114

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - sbt-scalatest

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent -5.12 56.55 $\color{green}{\blacktriangledown}$ -61.67 55.43 $\color{green}{\blacktriangledown}$ -60.55 56/210
agentEvpProxy -4.57 n/a n/a n/a n/a -

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - okhttp

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 20.19 19.59 $\color{red}{\blacktriangle}$ +0.60 19.99 $\color{red}{\blacktriangle}$ +0.20 34/108
agentless 18.80 18.82 $\color{green}{\blacktriangledown}$ -0.02 19.20 $\color{green}{\blacktriangledown}$ -0.40 30/105
agentlessCodeCoverage 21.73 22.54 $\color{green}{\blacktriangledown}$ -0.81 22.54 $\color{green}{\blacktriangledown}$ -0.81 30/106
agentlessLineCoverage 39.03 38.67 $\color{red}{\blacktriangle}$ +0.36 38.67 $\color{red}{\blacktriangle}$ +0.36 30/107

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - heliboard

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 9.08 9.92 $\color{green}{\blacktriangledown}$ -0.84 10.13 $\color{green}{\blacktriangledown}$ -1.05 31/107

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - reactive-streams-jvm

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 24.77 21.65 $\color{red}{\blacktriangle}$ +3.12 21.65 $\color{red}{\blacktriangle}$ +3.12 42/122
agentless 19.28 18.82 $\color{red}{\blacktriangle}$ +0.46 19.20 $\color{red}{\blacktriangle}$ +0.08 36/115
agentlessCodeCoverage 20.48 19.59 $\color{red}{\blacktriangle}$ +0.89 19.99 $\color{red}{\blacktriangle}$ +0.49 36/113
agentlessLineCoverage 27.01 26.45 $\color{red}{\blacktriangle}$ +0.56 26.45 $\color{red}{\blacktriangle}$ +0.56 22/59

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - netflix-zuul

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 30.57 87.80 $\color{green}{\blacktriangledown}$ -57.23 87.80 $\color{green}{\blacktriangledown}$ -57.23 35/112
agentless 23.26 81.05 $\color{green}{\blacktriangledown}$ -57.79 81.05 $\color{green}{\blacktriangledown}$ -57.79 32/109
agentlessCodeCoverage 31.16 95.12 $\color{green}{\blacktriangledown}$ -63.96 95.12 $\color{green}{\blacktriangledown}$ -63.96 32/109
agentlessLineCoverage 40.63 111.62 $\color{green}{\blacktriangledown}$ -70.99 111.62 $\color{green}{\blacktriangledown}$ -70.99 34/114

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - sonar-kotlin

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent -30.68 12.87 $\color{green}{\blacktriangledown}$ -43.55 12.87 $\color{green}{\blacktriangledown}$ -43.55 41/117
agentless -31.87 11.65 $\color{green}{\blacktriangledown}$ -43.52 11.88 $\color{green}{\blacktriangledown}$ -43.75 34/110
agentlessCodeCoverage -29.39 15.11 $\color{green}{\blacktriangledown}$ -44.50 15.11 $\color{green}{\blacktriangledown}$ -44.50 34/110
agentlessLineCoverage -28.64 17.03 $\color{green}{\blacktriangledown}$ -45.67 17.38 $\color{green}{\blacktriangledown}$ -46.02 34/110

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@gnufede
gnufede marked this pull request as ready for review September 18, 2026 13:53
@gnufede
gnufede requested review from a team as code owners September 18, 2026 13:53
@gnufede
gnufede requested review from AlexeyKuznetsov-DD and removed request for a team September 18, 2026 13:53
@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - jolokia

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 94.49 93.23 $\color{red}{\blacktriangle}$ +1.26 95.12 $\color{green}{\blacktriangledown}$ -0.63 36/119
agentless 88.87 89.58 $\color{green}{\blacktriangledown}$ -0.71 89.58 $\color{green}{\blacktriangledown}$ -0.71 37/118
agentlessCodeCoverage 98.03 99.00 $\color{green}{\blacktriangledown}$ -0.97 99.00 $\color{green}{\blacktriangledown}$ -0.97 38/118
agentlessLineCoverage 99.00 101.00 $\color{green}{\blacktriangledown}$ -2.00 101.00 $\color{green}{\blacktriangledown}$ -2.00 35/120

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@datadog-official datadog-official Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Datadog Autotest: PASS

More details

The dynamic retry path selects one duration bucket from the first test attempt. It keeps the selected execution limit for later attempts.

Was this helpful? React 👍 or 👎

Open Bits AI session

🤖 Datadog Autotest · Commit 63303ba · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - spring_boot

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 17.46 16.04 $\color{red}{\blacktriangle}$ +1.42 16.36 $\color{red}{\blacktriangle}$ +1.10 34/109
agentless 9.32 9.92 $\color{green}{\blacktriangledown}$ -0.60 9.92 $\color{green}{\blacktriangledown}$ -0.60 32/107
agentlessCodeCoverage 13.29 13.40 $\color{green}{\blacktriangledown}$ -0.11 13.40 $\color{green}{\blacktriangledown}$ -0.11 32/107
agentlessLineCoverage 23.16 22.54 $\color{red}{\blacktriangle}$ +0.62 22.09 $\color{red}{\blacktriangle}$ +1.07 33/107

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - sonar-java

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 15.46 15.41 $\color{red}{\blacktriangle}$ +0.05 12.87 $\color{red}{\blacktriangle}$ +2.59 37/119
agentless -17.08 10.97 $\color{green}{\blacktriangledown}$ -28.05 10.97 $\color{green}{\blacktriangledown}$ -28.05 37/119
agentlessCodeCoverage 44.04 77.88 $\color{green}{\blacktriangledown}$ -33.84 76.33 $\color{green}{\blacktriangledown}$ -32.29 37/119
agentlessLineCoverage 66.39 123.36 $\color{green}{\blacktriangledown}$ -56.97 116.18 $\color{green}{\blacktriangledown}$ -49.79 37/117

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - pass4s

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent -9.35 9.73 $\color{green}{\blacktriangledown}$ -19.08 9.73 $\color{green}{\blacktriangledown}$ -19.08 27/101
agentless -10.13 8.63 $\color{green}{\blacktriangledown}$ -18.76 9.54 $\color{green}{\blacktriangledown}$ -19.67 27/101
agentlessCodeCoverage -5.15 15.41 $\color{green}{\blacktriangledown}$ -20.56 15.72 $\color{green}{\blacktriangledown}$ -20.87 25/95

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@amarziali amarziali left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed locally at head 63303ba. No actionable findings.

@daniel-mohedano

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-09-23 14:20:45 UTC ℹ️ Start processing command /merge


2026-09-23 14:20:50 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-09-23 15:36:46 UTC ℹ️ MergeQueue: This merge request was merged

@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot merged commit 219b24c into master Sep 23, 2026
816 of 818 checks passed
@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot deleted the gnufede/dynamic-atr-retries branch September 23, 2026 15:36
@github-actions github-actions Bot added this to the 1.67.0 milestone Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: ci visibility Continuous Integration Visibility tag: ai generated Largely based on code generated by an AI or LLM type: feature Enhancements and improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants