Skip to content

Fix ATR eligibility with fail-fast test ordering - #12508

Merged
gh-worker-dd-mergequeue-cf854d[bot] merged 2 commits into
masterfrom
juan-fernandez/fix-atr-fail-fast-eligibility
Sep 16, 2026
Merged

gh-worker-dd-mergequeue-cf854d[bot] merged 2 commits into
masterfrom
juan-fernandez/fix-atr-fail-fast-eligibility

Conversation

@juan-fernandez

@juan-fernandez juan-fernandez commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

What Does This Do

Fixes an interaction between Auto Test Retries (ATR) and fail-fast test ordering.

Fail-fast test ordering loads the known-flaky test list so those tests can run first. ExecutionStrategy incorrectly treated that list's availability as "retry only known flakes," even when DD_CIVISIBILITY_FLAKY_RETRY_ONLY_KNOWN_FLAKES=false. Consequently, known-flaky failures retried while ordinary failures did not.

After this change, membership in the known-flaky list affects ATR eligibility only when known-flakes-only retries are explicitly enabled.

Motivation

With ATR and fail-fast ordering enabled:

  • Expected: Any otherwise-eligible failed test can be retried. The known-flaky list is used only to determine test execution order.
  • Actual: Only failed tests present in the known-flaky list were retried.

Loading data for one feature should not implicitly enable a separate ATR restriction.

Additional Notes

The JUnit 5 regression covers fail-fast ordering with an empty known-flaky list while known-flakes-only retries remain disabled. It fails without the ExecutionStrategy change because no retry attempts are emitted.

Validated with:

  • :dd-java-agent:instrumentation:junit:junit-5:junit-5.3:test --tests JUnit5Test
  • :dd-java-agent:agent-ci-visibility:test --tests datadog.trace.civisibility.test.ExecutionStrategyTest
  • :dd-java-agent:agent-ci-visibility:spotlessCheck
  • :dd-java-agent:instrumentation:junit:junit-5:junit-5.3:spotlessCheck

Contributor Checklist

Jira ticket: N/A

@juan-fernandez juan-fernandez added type: bug fix Bug fix comp: ci visibility Continuous Integration Visibility labels Sep 15, 2026
@juan-fernandez
juan-fernandez marked this pull request as ready for review September 15, 2026 14:41
@juan-fernandez
juan-fernandez requested review from a team as code owners September 15, 2026 14:41
@juan-fernandez
juan-fernandez requested review from ValentinZakharov and removed request for a team September 15, 2026 14:41
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-15T14:45:02.757882Z 1707543 Draft marked ready
🔒 Security Review ✅ Completed 2026-09-15T14:44:39.981788Z 1707543 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 17075435ad

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@datadog-datadog-prod-us1-2 datadog-datadog-prod-us1-2 Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Datadog Autotest: PASS

More details

The retry condition now checks the explicit known-flakes-only setting before it checks flaky-list membership. Fail-fast ordering can load an empty flaky list without stopping retries for other eligible tests.

Was this helpful? React 👍 or 👎

Open Bits AI session

🤖 Datadog Autotest · Commit 1707543 · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

@datadog-datadog-prod-us1-2

This comment has been minimized.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - nebula-release-plugin

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 37.57 36.42 $\color{red}{\blacktriangle}$ +1.15 36.42 $\color{red}{\blacktriangle}$ +1.15 34/110
agentless 36.52 35.70 $\color{red}{\blacktriangle}$ +0.82 35.70 $\color{red}{\blacktriangle}$ +0.82 34/110
agentlessCodeCoverage 44.38 44.48 $\color{green}{\blacktriangledown}$ -0.10 44.48 $\color{green}{\blacktriangledown}$ -0.10 34/110
agentlessLineCoverage 56.92 55.43 $\color{red}{\blacktriangle}$ +1.49 55.43 $\color{red}{\blacktriangle}$ +1.49 34/110

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - sbt-scalatest

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 55.20 55.43 $\color{green}{\blacktriangledown}$ -0.23 55.43 $\color{green}{\blacktriangledown}$ -0.23 56/204
agentEvpProxy 56.22 n/a n/a n/a n/a -

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - reactive-streams-jvm

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 20.85 21.65 $\color{green}{\blacktriangledown}$ -0.80 21.65 $\color{green}{\blacktriangledown}$ -0.80 42/119
agentless 19.05 19.20 $\color{green}{\blacktriangledown}$ -0.15 19.20 $\color{green}{\blacktriangledown}$ -0.15 36/112
agentlessCodeCoverage 19.81 19.59 $\color{red}{\blacktriangle}$ +0.22 19.99 $\color{green}{\blacktriangledown}$ -0.18 36/110
agentlessLineCoverage 26.88 26.45 $\color{red}{\blacktriangle}$ +0.43 26.98 $\color{green}{\blacktriangledown}$ -0.10 19/59

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - netflix-zuul

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 88.01 87.80 $\color{red}{\blacktriangle}$ +0.21 87.80 $\color{red}{\blacktriangle}$ +0.21 34/108
agentless 81.94 81.05 $\color{red}{\blacktriangle}$ +0.89 81.05 $\color{red}{\blacktriangle}$ +0.89 32/106
agentlessCodeCoverage 96.54 95.12 $\color{red}{\blacktriangle}$ +1.42 95.12 $\color{red}{\blacktriangle}$ +1.42 32/106
agentlessLineCoverage 111.58 111.62 $\color{green}{\blacktriangledown}$ -0.04 111.62 $\color{green}{\blacktriangledown}$ -0.04 32/109

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - pass4s

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 10.12 9.73 $\color{red}{\blacktriangle}$ +0.39 9.92 $\color{red}{\blacktriangle}$ +0.20 28/97
agentless 9.17 7.65 $\color{red}{\blacktriangle}$ +1.52 9.35 $\color{green}{\blacktriangledown}$ -0.18 28/97
agentlessCodeCoverage 13.27 16.36 $\color{green}{\blacktriangledown}$ -3.09 15.72 $\color{green}{\blacktriangledown}$ -2.45 25/91

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - heliboard

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 11.54 10.33 $\color{red}{\blacktriangle}$ +1.21 10.13 $\color{red}{\blacktriangle}$ +1.41 31/104

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@dd-octo-sts

dd-octo-sts Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

🟡 Java Benchmark SLOs — Performance SLO warning (near threshold)

Suite Status
Startup 🟡 warning

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 14.03 s 14.06 s [-1.1%; +0.6%] (no difference)
startup:insecure-bank:tracing:Agent 12.95 s 12.98 s [-1.0%; +0.6%] (no difference)
startup:petclinic:appsec:Agent 17.54 s 17.35 s [+0.2%; +1.9%] (maybe worse)
startup:petclinic:iast:Agent 17.44 s 17.56 s [-1.5%; +0.2%] (no difference)
startup:petclinic:profiling:Agent 17.42 s 17.37 s [-0.9%; +1.4%] (no difference)
startup:petclinic:sca:Agent 17.58 s 17.42 s [-0.1%; +1.8%] (no difference)
startup:petclinic:tracing:Agent 16.57 s 16.52 s [-0.8%; +1.4%] (no difference)

Commit: bb92d653 · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - sonar-kotlin

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 13.12 12.87 $\color{red}{\blacktriangle}$ +0.25 12.87 $\color{red}{\blacktriangle}$ +0.25 41/113
agentless 13.28 11.65 $\color{red}{\blacktriangle}$ +1.63 11.88 $\color{red}{\blacktriangle}$ +1.40 34/106
agentlessCodeCoverage 15.70 14.81 $\color{red}{\blacktriangle}$ +0.89 15.11 $\color{red}{\blacktriangle}$ +0.59 34/106
agentlessLineCoverage 16.40 17.38 $\color{green}{\blacktriangledown}$ -0.98 17.73 $\color{green}{\blacktriangledown}$ -1.33 34/106

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - jolokia

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 93.46 93.23 $\color{red}{\blacktriangle}$ +0.23 95.12 $\color{green}{\blacktriangledown}$ -1.66 37/116
agentless 89.65 89.58 $\color{red}{\blacktriangle}$ +0.07 89.58 $\color{red}{\blacktriangle}$ +0.07 36/113
agentlessCodeCoverage 98.15 99.00 $\color{green}{\blacktriangledown}$ -0.85 99.00 $\color{green}{\blacktriangledown}$ -0.85 37/113
agentlessLineCoverage 100.88 101.00 $\color{green}{\blacktriangledown}$ -0.12 101.00 $\color{green}{\blacktriangledown}$ -0.12 37/116

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - okhttp

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 19.45 19.59 $\color{green}{\blacktriangledown}$ -0.14 19.99 $\color{green}{\blacktriangledown}$ -0.54 33/105
agentless 18.40 19.20 $\color{green}{\blacktriangledown}$ -0.80 19.20 $\color{green}{\blacktriangledown}$ -0.80 31/102
agentlessCodeCoverage 22.39 22.54 $\color{green}{\blacktriangledown}$ -0.15 22.54 $\color{green}{\blacktriangledown}$ -0.15 31/103
agentlessLineCoverage 37.58 38.67 $\color{green}{\blacktriangledown}$ -1.09 38.67 $\color{green}{\blacktriangledown}$ -1.09 31/104

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - spring_boot

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 16.72 16.04 $\color{red}{\blacktriangle}$ +0.68 16.36 $\color{red}{\blacktriangle}$ +0.36 34/105
agentless 9.63 9.92 $\color{green}{\blacktriangledown}$ -0.29 9.92 $\color{green}{\blacktriangledown}$ -0.29 32/103
agentlessCodeCoverage 13.10 13.40 $\color{green}{\blacktriangledown}$ -0.30 13.40 $\color{green}{\blacktriangledown}$ -0.30 32/103
agentlessLineCoverage 22.15 22.54 $\color{green}{\blacktriangledown}$ -0.39 22.09 $\color{red}{\blacktriangle}$ +0.06 32/102

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

CI Visibility Test Environment - sonar-java

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 49.72 11.65 $\color{red}{\blacktriangle}$ +38.07 11.65 $\color{red}{\blacktriangle}$ +38.07 37/116
agentless 42.86 8.98 $\color{red}{\blacktriangle}$ +33.88 11.42 $\color{red}{\blacktriangle}$ +31.44 37/116
agentlessCodeCoverage 76.86 74.82 $\color{red}{\blacktriangle}$ +2.04 74.82 $\color{red}{\blacktriangle}$ +2.04 37/116
agentlessLineCoverage 162.44 111.62 $\color{red}{\blacktriangle}$ +50.82 111.62 $\color{red}{\blacktriangle}$ +50.82 36/113

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@daniel-mohedano

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-09-16 12:38:44 UTC ℹ️ Start processing command /merge


2026-09-16 12:38:50 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-09-16 13:52:57 UTC ℹ️ MergeQueue: This merge request was merged

@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot merged commit c595628 into master Sep 16, 2026
601 checks passed
@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot deleted the juan-fernandez/fix-atr-fail-fast-eligibility branch September 16, 2026 13:52
@github-actions github-actions Bot added this to the 1.67.0 milestone Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: ci visibility Continuous Integration Visibility type: bug fix Bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants