Repository navigation
Implement dynamic auto test retries - #12467
Conversation
Add dynamic Auto Test Retries (ATR) budgets based on test duration, instead of the flat per-test retry limit. When enabled, the number of retries allowed for a test is determined by the duration of its initial attempt, using the same duration buckets as Early Flake Detection (5s / 10s / 30s / 5m / >5m). Two new env vars control the behavior: - DD_CIVISIBILITY_DYNAMIC_ATR_ENABLED — enables duration-based ATR budgets. When unset/false, ATR stays on its existing flat-limit path. - DD_CIVISIBILITY_DYNAMIC_ATR_BUCKETS — optionally overrides the five duration-based retry budgets with five positive comma-separated integers in [1, 20]. When unset/empty, the EFD retry settings from the backend are used. The new handler classifies each test once (by its initial-attempt duration) and caches the resulting max-retries count for the lifetime of that test. The duration bucket index is computed via shared helpers on the EFD settings type, so bucket boundaries stay consistent with EFD (the existing EFD retry handler is refactored to use them too). Telemetry: records a `dynamic_atr_retries.enabled` count metric with a `has_custom_buckets` tag when dynamic ATR is enabled. Feature parity with dd-trace-py PR #20028.
This comment has been minimized.
This comment has been minimized.
🟢 Java Benchmark SLOs — All performance SLOs passed
PR vs. master results
Commit: Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion. |
CI Visibility Test Environment - nebula-release-pluginJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - sbt-scalatestJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - okhttpJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - heliboardJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - reactive-streams-jvmJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - netflix-zuulJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - sonar-kotlinJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - jolokiaJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
There was a problem hiding this comment.
More details
The dynamic retry path selects one duration bucket from the first test attempt. It keeps the selected execution limit for later attempts.
🤖 Datadog Autotest · Commit 63303ba · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest
CI Visibility Test Environment - spring_bootJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - sonar-javaJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
CI Visibility Test Environment - pass4sJob Status: 🟢 success
Baseline: median of Tests run to verify that CI Visibility behavior has not regressed in the current PR. |
|
/merge |
|
View all feedbacks in Devflow UI.
The expected merge time in
|
219b24c
into
master
Motivation
Add dynamic Auto Test Retries (ATR) budgets based on test duration, instead of the flat per-test retry limit. When enabled, the number of retries allowed for a test is determined by the duration of its initial attempt, using the same duration buckets as Early Flake Detection (5s / 10s / 30s / 5m / >5m).
Two new env vars control the behavior:
The new handler classifies each test once (by its initial-attempt duration) and caches the resulting max-retries count for the lifetime of that test. The duration bucket index is computed via shared helpers on the EFD settings type, so bucket boundaries stay consistent with EFD (the existing EFD retry handler is refactored to use them too).
Telemetry: records a
dynamic_atr_retries.enabledcount metric with ahas_custom_bucketstag when dynamic ATR is enabled.Feature parity with dd-trace-py PR #20028.