Skip to content

publish_serve_cycle intermittently hangs to the 120s nextest timeout #336

Description

@LKSNDRTMLKV

dpp-vault::publish_serve_cycle::published_passport_is_served_as_the_payload_its_proof_signed timed out at nextest's 120-second ceiling in run 35048585684, failing the integration tier. 380 of 381 tests passed; this one was the whole failure.

Re-running the same job on the same commit passed. The commit under test (#321) touches only crates/dpp-seal/src/trustlist/, which has no runtime caller — so nothing in the change could reach this test's code path.

It is a hang, not a slow test

Worth separating, because the two want different fixes.

scripts/slow-test-check.sh runs with a 10-second budget and this test is not on the allowlist. It does not appear in the slow-test report on any passing run, so on a normal run it completes inside 10 seconds. On the failing run it was still going at 120.

That is not a test drifting over budget. It is a test that occasionally does not finish at all — twelve times its own normal ceiling, then killed.

The slow-test gate did its job and reported it:

ERROR: tests over the 10s budget:
  120.014s  dpp-vault::publish_serve_cycle::published_passport_is_served_as_the_payload_its_proof_signed

…but that message reads as "make the test faster", which is the wrong instruction here. The fix is to find what it waits on forever.

Where to look

The test is #[tokio::test(flavor = "multi_thread")] and its first three statements each stand something up:

let pg = start_postgres().await;
let vault_url = start_vault(pg.dal.clone()).await;
seed_complete_operator(&pg.dal).await;

then drives the real HTTP surface through TestClient. Candidates, roughly in order of how much they would explain a 120s hang against a 10s norm:

  1. start_postgres / start_vault — container or bind contention under a loaded runner. The allowlist's migration_0024 entry notes that start_pg_before boots a server that cannot use the shared template, which implies the shared template is what makes the normal path fast; a fallback or a race around that would fit the timing gap exactly.
  2. An HTTP call with no client-side timeout. If TestClient has none, any request that never gets a response parks until nextest kills the process, which is precisely the observed shape — no assertion failure, no panic, just silence.
  3. Port or pool exhaustion across the 381 tests in the tier, biting whichever test happens to run last. This one was 381/381.

(2) is worth checking first whatever the cause turns out to be: a test client without a timeout converts every backend stall into a 120-second CI failure with no diagnostic, which is why this run says nothing about what it was waiting for.

Why it is worth fixing rather than re-running

The integration tier is a required status check and merges are strict-up-to-date, so every PR has to pass it and a flake here costs a full ~11-minute cycle each time it fires. It also trains the reflex of re-running a red gate, which is the reflex that eventually merges a real failure.

Found while merging #321; the failure was confirmed unrelated before the re-run rather than after — main was green at 9efa7a4 and this was the only failure in the previous 40 runs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    type/defectSomething published or encoded here is wrong or unbackable nowurgency/nextBlocks work already scheduled

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions