Skip to content

fix(seal): read the seal mode from the backend, not a literal - #287

Merged
LKSNDRTMLKV merged 1 commit into
mainfrom
fix/seal-mode-from-capabilities
Sep 10, 2026
Merged

LKSNDRTMLKV merged 1 commit into
mainfrom
fix/seal-mode-from-capabilities

Conversation

@LKSNDRTMLKV

Copy link
Copy Markdown
Member

Found by a runtime pass against a stack rebuilt from main: SEAL_PROVIDER=local could not produce a single seal, under any configuration.

The bug

seal_drain.rs named SealMode::ProviderSeal as a literal in the request it builds for every row.

That is not a setting. Core's SealPort states the rule: "the mode decides whose attestation the seal is." ProviderSeal says a qualified provider holds the key and attests on the operator's behalf. Only the QTSP backend advertises it.

The local development sealer signs under the operator's own self-signed certificate and advertises OperatorSeal — correctly. So the adapter's capability check refused every request the drain built for it, exactly as it should. Each row then backed off through eight attempts and reached exhausted.

Observed on a live node:

asked for:  format Cades, mode ProviderSeal,  level BaselineB, envelope Detached
advertises: formats [Cades], modes [OperatorSeal], levels [BaselineB], envelopes [Detached]

Meanwhile GET /vault/api/v1/seal reported sealingConfigured: true throughout, and nothing was logged until exhausted climbed — hours later.

A literal could never have been right for more than one backend.

The fix

mode now travels from the composition root, read off the backend's advertised capabilities — the same route credential and conformance_level already take, and for the same stated reason: it is a property of the provider arrangement, not of the drain loop.

Two consequences worth calling out:

  • A node that would enqueue seal rows it cannot drain now refuses to boot. Every input to that decision exists at startup. It was instead discovered per row, after publish, in silence. Gated on drains, so a node with sealing off is unaffected.
  • The failure names the axis that actually mismatched. The adapter's error listed all four dimensions and closed with "Set SEAL_CONFORMANCE_LEVEL to a level this backend supports" — wrong guidance whenever the level was not the problem. It cost me a restart chasing the wrong variable. Where the mode is the mismatch, the message now says so, and says it is not settable.

Still true, and deliberate

A node running SEAL_PROVIDER=local needs SEAL_CONFORMANCE_LEVEL=B. The local sealer advertises BaselineB only, and refusing to silently satisfy a request for a level that outlives its own certificate is the documented intent. The difference is that this is now a boot error naming the variable rather than eight silent retries.

Tests

Six new unit tests. The regression one is a_backend_that_signs_under_its_own_certificate_is_asked_for_an_operator_seal; a_provider_held_backend_is_asked_for_a_provider_seal pins that the fix does not swing the other way; a_level_the_backend_cannot_reach_fails_the_boot_and_names_the_remedy pins the message; a_node_that_does_not_seal_is_not_checked pins the drains gate.

just check green — 1077 tests. check-integration caught the feature-gated seal_outbox suite calling drain_once at the old arity, which is precisely what that step exists for; those call sites now pass ProviderSeal, matching the QTSP mock they drive.

End-to-end verification against a live node is in the PR thread.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant