Skip to content

feat(config): deliver gateway configuration over the supervisor session - #4320

Draft
pimlock wants to merge 36 commits into
mainfrom
feat/1731-gateway-config-push/pimlock
Draft

pimlock wants to merge 36 commits into
mainfrom
feat/1731-gateway-config-push/pimlock

Conversation

@pimlock

@pimlock pimlock commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Supervisors poll the gateway for configuration every 10 seconds. This PR adds an opt-in push path: with config_delivery_mode = "push", the gateway delivers configuration over the existing ConnectSupervisor session, and supervisors apply and acknowledge it. poll stays the default, and older supervisors and gateways keep polling.

Durable completion operations (wait_mode, GetConfigUpdateOperation) moved to a follow-up PR so this one stays reviewable. policy set --wait keeps confirming application through policy status, which push sessions report from apply results.

Related Issue

Part of #1731. Durable operations are deferred to a follow-up PR; the issue's acceptance criteria need a matching amendment.

This replaces the stacked PRs #3244, #3265 and #3273; their review history stays there.

Changes

Delivery

  • Bootstrap, update, result and admission contracts on ConnectSupervisor. Supervisors advertise supports_config_apply, and the gateway answers with config_apply_enabled, so older supervisors keep polling.
  • Each push session has its own delivery task. A publication marks the components the session must rebuild; the task takes the marks once it holds a build permit, so repeated publications coalesce and a build always reads state at least as new as the publications it covers.
  • Build permits come from two FIFO pools sized from the database pool. Fanout work uses the shared pool; a sandbox-scoped change takes a shared or reserved permit, whichever frees first, so fleet-wide updates never starve it. Concurrent fanouts are served in order rather than interleaved per workspace.
  • A failed build retries with per-session backoff from 1 to 30 seconds. Each session also rebuilds its sandbox configuration every 30 seconds, starting 30-60 seconds after it connects, to repair missed publications.
  • Reconciliation rebuilds only the sandbox configuration. When its provider_env_revision changed, the supervisor reports AWAITING_COMPONENT and the gateway builds the provider environment, the same rule polling supervisors follow. Reconciliation therefore never calls credential drivers for unchanged providers.
  • Provider changes rebuild only the sandboxes that attach the provider.
  • Updates are held until the supervisor reports its bootstrap result, so a publication that races a new session is dropped when it matches the bootstrap instead of being applied twice.
  • In HA, other replicas send the session owner a secret-free PeerNotifyConfigUpdate, and the owner rebuilds from shared state. Single-replica gateways skip the owner lookup.

Cheaper builds (also benefit poll mode)

  • Interceptor provider-profile sources are served from a cached snapshot, refreshed every provider_profile_source_refresh_interval_seconds (default 10). A changed revision is pushed to connected sessions. Reads fail closed before the first fetch and once refreshes have failed for three intervals.
  • Provider environment builds read each provider's refresh states once instead of twice, and load providers concurrently.
  • One build attempt loads the provider catalog, provider records, global settings, and latest sandbox policy once and builds both components from them. The bootstrap, combined live builds, and the provider readiness snapshot use this shared load. The bootstrap also checks that both components agree on the attachment epoch and effective policy hash, not only the provider revision. Live components keep their own deadlines, deliveries, and retries, and only the provider environment resolves credentials.

Streamed apply

  • In push mode, supervisors prepare the startup policy against the image, apply the bootstrap before starting the workload, apply and acknowledge live updates, and stop polling.
  • Polled and delivered configuration go through one supervisor apply path. Quarantine recovery, credential revocation and retention, OCSF events, and stale-generation checks behave the same in both modes.
  • A supervisor whose middleware is unreachable at startup starts degraded, as polling does.
  • Streamed admission ends image preparation, the same as a polled configuration report does.
  • The supervisor's session-established OCSF event reports config_delivery=push|poll.

Observability

  • New metrics, documented in the gateway metrics reference: build duration (histogram), dirty and waiting sessions, deliveries, apply results, bootstraps, peer notifications, and provider-profile source refreshes and snapshot age.

Testing

  • mise run pre-commit passes
  • cargo test -p openshell-server --features test-support, including Tokio paused-clock tests for coalescing, the sandbox reserve, lane raises, failure backoff, per-session reconcile and session close
  • Workspace unit tests. Failures outside this change: openshell-core e2fsprogs (ETXTBSY under parallel runs), openshell-isolation-interface network-namespace and openshell-supervisor-network proxy tests that time out on loopback on the test host. None of those crates changed.
  • mise run go:ci, mise run sdk:ts:ci
  • PostgreSQL suites against postgres:16-alpine (2/2; the other three were operation tests and moved with operations)
  • mise run e2e:rust:push: conformance 6/6, config_push 1/1, live_policy_update 4/4, policy_activation 1/1
  • mise run e2e:rust (poll): conformance, 116/116 Rust e2e tests

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO)
  • Existing polling APIs remain available

@pimlock pimlock added gator:approval-needed Gator completed review; maintainer approval needed gator:blocked Gator is blocked by process or repository gates test:e2e Requires end-to-end coverage test:e2e-kubernetes Requires Kubernetes end-to-end coverage labels Oct 7, 2026
@github-actions

This comment was marked as outdated.

@github-actions

This comment was marked as outdated.

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

@pimlock pimlock removed gator:approval-needed Gator completed review; maintainer approval needed gator:blocked Gator is blocked by process or repository gates labels Oct 7, 2026
@pimlock pimlock changed the title feat(config)!: deliver gateway configuration over the supervisor session feat(config): deliver gateway configuration over the supervisor session Oct 7, 2026
Add an opt-in push path for gateway-owned supervisor configuration,
selected by the gateway setting config_delivery_mode = "poll" | "push".
Poll stays the default, and polling keeps working for older supervisors
and gateways.

Delivery:
- Add bootstrap, snapshot, update, result and admission contracts on
  ConnectSupervisor. Supervisors advertise supports_config_snapshots and
  supports_config_apply; the gateway answers with config_apply_enabled.
- Configuration delivery is a coalescing DeliveryQueue drained by one
  on-demand delivery worker. Admission never drops updates, fanouts walk
  connected sandboxes as cursors, and sandbox-scoped work keeps reserved
  build slots.
- Provider changes rebuild only the sandboxes that attach the provider.
- Across replicas, other gateways send the session owner a secret-free
  PeerNotifyConfigUpdate and the owner rebuilds from shared state. An
  owner reconciler repairs missed notifications.

Streamed apply:
- Supervisors that enable apply prepare the startup policy against the
  image, apply the bootstrap before starting the workload, apply and
  acknowledge live updates, and stop polling.
- A failed live registry reload keeps the last-known-good registry.

Durable completion:
- Sandbox policy and settings updates commit desired state and a durable
  ConfigUpdateOperation atomically. Callers choose wait_mode COMMIT_ONLY
  or WAIT_FOR_COMPLETION, with idempotency keys and GetConfigUpdateOperation.
- Operations resolve as applied, degraded, failed, superseded, inactive or
  cancelled, and survive client timeouts and gateway restarts.

This combines the previously stacked changes from #3244, #3265 and #3273.

Refs #1731

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock
pimlock force-pushed the feat/1731-gateway-config-push/pimlock branch from b4bc6e9 to 0768489 Compare October 7, 2026 23:04
@pimlock
pimlock marked this pull request as draft October 7, 2026 23:19
@copy-pr-bot

This comment was marked as outdated.

Completion compares only the dimension an operation changed: the policy
version for policy updates, the settings revision for settings updates.
The per-sandbox fence and the in-transaction reads of the other
dimension only made an uncorrelated value consistent, so remove them.

Drop the sandbox_config_fences table from migration 009 and the fence
and target helpers from both persistence backends. Operations now store
the target the request set. Exercise all three operation write paths on
SQLite and PostgreSQL.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Remove SupervisorHello.supports_config_snapshots. supports_config_apply is
the only configuration capability; the gateway sends a bootstrap exactly
when it enables apply.

Report a component that waits for its matching generation with the new
CONFIG_APPLY_OUTCOME_AWAITING_COMPONENT instead of a FailedClosed result
carrying the generation_mismatch string. Operations skip it because it is
not terminal.

UpdateConfigRequest.wait_timeout is new, so it takes field 14 directly
instead of reserving a never-released scalar.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Add an e2e:rust:push lane that starts the Docker gateway with
config_delivery_mode = "push" and runs CLI conformance, the live policy
update and policy activation tests, and a new config_push test. The new
test requires policy set --wait to complete from the supervisor's
acknowledgement and report its operation, so the lane fails if delivery
silently falls back to polling.

OPENSHELL_E2E_DOCKER_TEST now accepts several space-separated targets.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
…h one path

Session-delivered configuration re-implemented the poll loop's reconcile
logic and had drifted from it. Both now run ConfigRuntime::reconcile, which
owns the applied-state trackers. Polling and session delivery differ only in
where snapshots and provider environments come from and in how results reach
the gateway: ReportSandboxConfiguration and ReportPolicyStatus for polling,
session acknowledgements for delivery.

This fixes delivered configuration that:
- reported Applied for a snapshot restoring the policy of a rejected update
  while the runtime stayed in its fail-closed quarantine;
- kept static provider credentials after a provider snapshot that could not
  be installed;
- kept rotating credentials for removed middleware services;
- skipped the policy reload events, stale generation checks, and policy
  activation failure readiness.

Delivered provider snapshots now use the polling conversion, so invalid
expiration times are rejected and unknown readiness reasons withhold
credentials. A startup whose bootstrap middleware is unreachable starts
degraded, as polling does, and retries the registry while idle.

Result and revision builders live in the session module. A bootstrap that
the runtime cannot apply answers every component as an update does instead
of sending an empty result. Updates on a session without streamed apply are
ignored rather than answered.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Drop the sandbox_config_fences table and the other-dimension target reads
it made consistent. Every fenced write runs under the sandbox sync guard,
and each operation reads only the target of the dimension its request
mutated.

Remove UpdateConfigRequest.idempotency_key and renumber wait_timeout to
13. request_id replay returns the committed operation, and a retried
WAIT_FOR_COMPLETION waits on it again. Operations are named by their ID
and no longer store response fields that only idempotent replay read.

Record no operations in poll mode. In push mode, mark completion of
operations for polling supervisors with an INACTIVE state and UNSUPPORTED
outcome instead of matching error text.

Delete a sandbox's operations with the sandbox, and refuse new operations
once deletion has begun so none can outlive it. Provider receipts are
kept for mutation replay.

Remove the projection repair, the legacy time-field mapping, and the
dimension fallback, which only unreleased builds could need. Remove the
handler-internal completion wait; the UpdateConfig wrapper owns it.
Settings and operation rows share column derivation and sandbox
projection logic across the SQLite and PostgreSQL backends.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
setPolicy({ wait: true }) returned immediately when the gateway reported
no completion operation, as in the default poll mode and with older
gateways. Restore the getConfig policy-hash poll for that case and keep
the durable-operation checks for gateways that track completion.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
…ition

An operation insert for a sandbox in the Deleting phase returned a
resource-version conflict, which callers read as "modified concurrently;
please retry". Retrying cannot succeed, so return FAILED_PRECONDITION with
"sandbox is being deleted" instead.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Push mode streams configuration only to supervisors that apply it. Remove the
optional shadow bootstrap and its short build timeout, the oversized-bootstrap
stripping, the unacknowledged delivery branch, and the delivery-worker path
that finished operations for snapshot-only sessions. A registration now always
follows a built bootstrap.

A session's mode carries its configuration slots, and only apply sessions
hold per-component delivery state, replacing the separate slots option and
acknowledgement flag.

Deliver through the session registry instead of a transport trait with one
implementation, and replace ConfigComponents::SANDBOX_AND_PROVIDER with the
identical ALL.

Convert the delivery tests that connected snapshot-only sessions to apply
sessions. The stalled-credential test now proves that a reconnect bootstrap
waiting on credentials leaves the live session serving relays.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
pimlock added 18 commits October 7, 2026 23:36
…_delivery

Move the streamed-apply session protocol out of supervisor_session.rs into
config_delivery/session.rs: the bootstrap build, the startup candidate and
prepared exchange, admission, per-component delivery state, and result
handling. Relays, peer routing, and session lifecycle stay in place, and the
registry exposes the configuration state of one exact session to the new
module.

One borrowed snapshot view now builds every revision and fingerprint, and a
single ExpectedBootstrap validates the component results and admission of
a bootstrap before anything is persisted. The session loop and message
handler take the accepted session as one value instead of twelve
arguments.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
A supervisor that cannot reach its middleware at startup enforces the policy
with built-in middleware and reports the sandbox component as Degraded. The
gateway recorded no policy load for that outcome, so the revision stayed
pending. Treat Degraded like Applied when recording policy status.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Sessions without streamed apply no longer receive configuration, so finishing
their untracked operations no longer republishes. Document when
UpdateConfigResponse omits its operation and what an inactive operation with
an unsupported outcome means.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
A supervisor whose bootstrap middleware was unreachable retried the
startup configuration on idle ticks even after a later delivery. Once
middleware recovered, the retry reinstalled the startup policy over a
later rejection's fail-closed quarantine, or over a later failed update,
and the gateway was never told. The retry now stops as soon as a
different configuration reconciles.

A reconnect that redelivered the degraded startup configuration while
middleware was still down answered FailedRetainedLastKnownGood with the
requested revision as the retained one. The gateway rejects that result
and ends the session, so every reconnect failed the same way. A failed
upgrade of the configuration the runtime already enforces now reports
Degraded with an accepted admission, and readiness stays bound to the
generation that enforces it instead of reporting a failed activation.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Deleting a sandbox removes its configuration operations, so a client waiting
on one received NOT_FOUND instead of the terminal cancellation the wait
contract promises. Report the operation the waiter last observed as
cancelled with "sandbox no longer exists", as reconciliation does.

The proposal merge path now reports an update to a deleting sandbox as a
failed precondition instead of an internal error.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Image preparation deadlines (#4038) leave preparation when a supervisor's
admission is recorded. ReportSandboxConfiguration did that, but admission
delivered over the supervisor session in push mode only stored the
admission, so every push-mode sandbox stayed Provisioning with
"Waiting for an authenticated supervisor configuration report" until its
deadline. Record the admission start, and any rejection, in the same CAS
that stores streamed admission.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
The CLI reports an unchanged policy without the operation the gateway
waited on, so assert the unchanged result and that the waited call
completes.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Port the SQLite and PostgreSQL check that settings, policy, and unchanged
request writes store only the dimension the operation changed, using the
single-dimension operation targets.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Record openshell_supervisor_config_build_duration_seconds per component
and outcome as a bucketed histogram, and add debug spans around snapshot
builds and credential resolution so a trace shows where a slow build
spends its time.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Durable sandbox configuration operations move to a follow-up change so
this one only delivers configuration over the supervisor session. Remove
UpdateConfigRequest.wait_mode and wait_timeout, UpdateConfigResponse.operation,
GetConfigUpdateOperation, the operation reconciler, migration 009, and the
CLI, Go, and TypeScript handling of operations. Provider receipts keep the
existing operation module from main.

Sandbox-scoped updates now publish their components directly after
commit. `--wait` keeps confirming application through policy status, which
push sessions report from apply results.

The push e2e lane proves push mode from the supervisor's bootstrap log line
instead of an operation ID.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
A single-replica gateway owns every supervisor session, so the post-build
owner index read is a wasted store query on every delivery. The session
check in deliver_config still refuses a replaced session.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
The 30-second owner reconciler rebuilt both components for every session,
so each pass called the credential drivers for every attached provider.
Polling supervisors fetch the provider environment only when the sandbox
configuration's provider_env_revision changes. Push now follows the same
rule: the reconciler rebuilds the sandbox configuration, and a changed
revision makes the supervisor wait for the provider environment, which the
gateway then builds through the existing counterpart path.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Every configuration build and every 10-second GetSandboxConfig poll called
the interceptor's SnapshotProviderProfiles RPC, which takes no arguments
and returns the same global snapshot. Serve interceptor sources from a
cached snapshot instead. The gateway fetches it before serving and then
every provider_profile_source_refresh_interval_seconds (default 10), and
publishes configuration to connected sessions when the revision changes.

Reads still fail closed: before the first successful fetch, and once
refreshes have failed for three intervals, they return UNAVAILABLE as a
failed fetch did before. Refresh outcomes and snapshot age are exported as
openshell_server_provider_profile_source_refreshes_total and
openshell_server_provider_profile_source_snapshot_age_seconds.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Provider environment builds load every attached provider's record and
refresh states, then validated key ownership by listing the refresh states
again for each provider. Validate against the states already loaded, and
load the providers concurrently. The resolvers no longer take a store.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Replace the global DeliveryQueue and its delivery worker with one task per
push session. A publication marks the components the session must rebuild
and the highest lane that asked; the task takes the marks once it holds a
build permit, so a build always reads state at least as new as every
publication it covers, and repeated publications coalesce into one build.

Build permits come from two FIFO semaphores sized from the database pool.
Fanout work uses the shared pool; sandbox-scoped work takes a shared or
reserved permit, whichever frees first, and a sandbox edit moves a session
already waiting for a fanout permit into that lane. Fanouts no longer
interleave per workspace; waiters are served in order.

A failed build now retries on its own with per-session backoff from 1 to
30 seconds instead of waiting for the next reconcile pass. The global
30-second reconciler becomes a per-session timer with a random offset.
Peer notifications keep their per-target coalescing and concurrency bound.

The fanout pass, active fanout, build skip, and per-lane pending metrics are
replaced by openshell_supervisor_config_dirty_sessions and
openshell_supervisor_config_waiting_sessions.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
- Convert a delivered sandbox configuration snapshot through the polling
  projection, as provider environment snapshots already are, so both paths
  share defaults and validation. The conversion lists every response field,
  so a new one cannot be missed.
- Drop the per-component sequence watermarks. The gateway sends at most one
  update per component in order and replaces unsent ones, so the stream
  never carries an older sequence.
- Test that the deferred loopback connector holds relay connects until the
  boundary connector is installed.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Describe the push-mode configuration delivery metrics and the interceptor
provider-profile source refresh metrics, including their label values.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
A publication that raced a session's bootstrap rebuilt both components and
sent them before the supervisor reported its bootstrap result, so the
gateway could not tell they matched what the supervisor already had. In
e2e runs a third of sandbox configuration and provider environment
deliveries were applied as IGNORED_DUPLICATE.

Treat the bootstrap like an update in flight: a rebuild that arrives first
waits as the pending snapshot, and is dropped when the bootstrap result
acknowledges the same snapshot or sent otherwise. The hold is bounded by the
same acknowledgement timeout as live updates.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@pimlock
pimlock force-pushed the feat/1731-gateway-config-push/pimlock branch from b9972d5 to f288455 Compare October 8, 2026 22:47
@copy-pr-bot

This comment was marked as outdated.

@pimlock

pimlock commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

/ok to test f288455

A replica without a peer endpoint records `local://<replica>` as its owner
endpoint. Relay routing rejected that placeholder, while configuration
peer notifications instead allowed only http:// and https:// endpoints.
Move the rule next to the owner record as `OwnerRecord::peer_endpoint`,
which treats empty and `local://` endpoints as not dialable, and use it for
relays, session redirects, and both kinds of configuration notification.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Supervisor sessions had three modes: Legacy, ReportsReadiness, and
ConfigApply. Legacy and ReportsReadiness both poll; they differed only in
when the gateway marks the sandbox ready, inferred from the supervisor
version. Supervisors from before streamed apply connect once their workload
runs and are ready on accept; newer ones connect before the workload starts
and report SupervisorRuntimeReady.

Report that fact directly instead: SupervisorHello.workload_pending is set
while the supervisor's workload has not started. Supervisors that predate
it leave it unset, which matches how they connect. SessionMode is now
Poll { ready_on_accept } or Push. A newer supervisor that reconnects to a
polling gateway after its workload started is now ready on accept, as on
main, and a reconnect during startup still waits for its readiness report.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
A polling gateway now marks a session ready on accept when its hello
reports the workload as running, so repeating SupervisorRuntimeReady on
such a reconnect only rewrote the same state. Send it only when the hello
reported the workload as pending and it has started since, which is the
case where the gateway is still waiting for it.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
A provider environment the supervisor failed to apply was never retried:
reconciliation rebuilt only the sandbox configuration, which the gateway
skips once acknowledged. Each session now tracks which components its
supervisor last failed to apply, and the 30-second reconcile rebuilds
them along with the sandbox configuration, matching poll mode's retry of
an unavailable provider environment.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Load the provider catalog, provider records, global settings, and latest
sandbox policy once per build attempt, and build the sandbox configuration
and provider environment from that captured view. Sandbox settings and
global policy version stay component-specific, and only the provider
environment resolves credentials.

The streamed bootstrap now loads inputs once per attempt and checks the
provider attachment epoch and effective policy hash as well as the
provider revision. Live delivery shares one load across the selected
components while keeping each component's deadline, delivery, and
failure independent. The provider readiness snapshot also builds both
components from one load.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Drop the unreleased provider_profile_source_refresh_interval_seconds
gateway setting. Interceptor provider-profile sources refresh every 10
seconds, matching the supervisor configuration poll interval.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Move streamed apply, the configuration runtime, the policy poll loop,
and settings helpers from lib.rs into config_runtime.rs, together with
the tests that exercise them. The code moves unchanged apart from
imports and visibility.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

test:e2e Requires end-to-end coverage test:e2e-kubernetes Requires Kubernetes end-to-end coverage

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant