fix(queue): a Dependent-suites no-verdict is infrastructure (steward re-queues) + the freeze RCA - #5801
Conversation
…IFO on the shared dind label Three queue entries in a row failed only on `Dependent suites (MeshWeaver.Plugins)`: the waiter hit its 45-min cap with NO verdict, while the Plugins candidate legs (5-14 min of work) waited 20-50 min for an `aks-silos-dind` runner behind PR and bake work. Records the measurement, the root cause (no order between gate work and PR work on one FIFO label at a hardware-bound cap), the fix in the repos that own it (Memex#587 gate lane with a higher PriorityClass; Plugins#2444 legs on that lane + stand down once the requesting core run has finished), and that the steward correctly REJECTS such an entry rather than re-queueing it into the same line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Correct the waiter deadlines, utilization definition, and steward timeout-versus-failure behavior.
Review effort: Lite
Findings: 1
What changed in this PR
Docs-only update documenting the 2026-09-27 merge-queue freeze, its runner-starvation cause, and remediation.
Changes:
- Records incident measurements and affected runs.
- Documents the dedicated gate lane and stand-down behavior.
- Clarifies steward handling and queue-size controls.
| File | Description |
|---|---|
src/MeshWeaver.Documentation/Data/Architecture/MergeQueue.md |
Adds the incident analysis, measurements, fixes, and operational guidance. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
|
||
| The paragraph above predicted it, and on 2026-09-27 09:00–11:00Z it happened: nothing merged for two | ||
| hours. Three entries in a row (core runs 36307979031, 36308367885, 36311194922) failed on ONE job, | ||
| `Dependent suites (MeshWeaver.Plugins)`. Each time its waiter hit the 45-minute cap **with no verdict**. |
There was a problem hiding this comment.
Agreed. Fixed in the next push: the text now distinguishes the waiter's 42-minute verdict deadline (await-dependent-verdict.py --deadline-minutes 42) from the job's 45-minute timeout-minutes.
…45-min cap (review) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ward re-queues it, capped Five core PRs (#5789 #5791 #5792 #5793 #5795) sat queue-rejected on 2026-09-27 for one reason: 'MeshWeaver.Plugins did not answer within 42 min (no verdict yet)' — the dependent's legs never got a runner, while every candidate that did run was green. The steward classed that as a gate/build failure ('never a flake') and left them out. - await-dependent-verdict.py: silence exits EXIT_NO_VERDICT=3 (still red, never a pass); every verdict-shaped red stays exit 1. - dotnet-test.yml: the wait step hands exit 3 to its own failing step, 'No verdict in time: the dependent's suites did not report (infrastructure)'. - merge-queue-steward.py: a Dependent-suites job that failed ONLY on that step is starved, not failed — re-queue kind=infra, capped 2 per head sha with the marker comment; beside a build failure or an uncatalogued assertion it still rejects; a verdict red still rejects. Nine new self-test rows (66 ok). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
ℹ️ Merge-queue steward: no action — removed from the queue with reason |
…a; steward table row Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

What
1. A Dependent-suites NO-VERDICT is infrastructure, and the steward re-queues it (capped)
Five core PRs (#5789, #5791, #5792, #5793, #5795) were
queue-rejectedon 2026-09-27 for one reason only:MeshWeaver.Plugins did not answer within 42 min (no verdict yet). The dependent's candidate legs never got a runner. Every candidate that did run that morning was green. The steward treated that job as a gate/build failure ("never a flake") and left the PRs out.await-dependent-verdict.py: silence now exits withEXIT_NO_VERDICT = 3. It is still red and never a pass. Every red that comes from a verdict still exits 1.dotnet-test.yml: the wait step passes exit 3 to its own failing step, No verdict in time: the dependent's suites did not report (infrastructure).merge-queue-steward.py: aDependent suites (MeshWeaver.Plugins)job that failed only on that step counts as starved, not failed.infra, capped at 2 per head sha, with the marker comment.actionlintis clean.The five PRs already rejected will not be re-queued by this change, because the steward acts only on new
dequeuedevents. They need a human re-queue once the queue flows.2.
MergeQueue.md(+ the waiter row inCrossRepoPairGate.md): the RCA of the freezeDependent suitesalone, with no verdict, when the waiter hit its 42-min deadline (inside the job's 45-min cap).aks-silos-dindwith PR and satellite work, and that queue is served FIFO. The legs ran 5–14 min but waited 20–50 min. The set was at a cap that the hardware quota imposes.aks-silos-dind-gatewith PriorityClassarc-runner-gate, which never preempts. It is applied and live, with max 12. Systemorph/MeshWeaver.Plugins#2444 moves the legs onto that lane and makes a leg stand down once its core run has finished.Pairs-with: none — no public surface changes (workflow, scripts and docs only).
🤖 Generated with Claude Code