Skip to content

bug(supervisor-network): boundary reconnect is followed by proxy exit and ControlSupervisorExited #3396

Description

@slopp

User Story

As an operator running a long-lived agent sandbox through OpenShell's RFC 0012 split supervisor/workload architecture, I need a transient Sandbox Protocol boundary failure to recover cleanly—or be rejected as an incompatible deployment before launch—so that an otherwise healthy workload is not terminated and left in Error.

Problem Statement

A Docker-backed sandbox ran successfully for approximately 15 minutes, including allowed proxied HTTPS traffic and successful in-place gateway/Sandbox Protocol credential renewal. The network-mediation proxy then exited with boundary unavailable.

The workload boundary detected the supervisor connection loss, froze the workload, and reported recovery about 4 ms later. Ten seconds after that apparent recovery, the boundary gRPC exchange failed. The supervisor stream ended, the gateway reported an empty Docker wait error, and the sandbox transitioned to Error with reason ControlSupervisorExited. The Docker driver then stopped the workload; the workload container itself was not OOM-killed and exited 0.

Repeated ReportEndpointStatus calls were already returning gRPC status 9 before the fatal event, while GetSandboxConfig and GetSandboxProviderEnvironment continued succeeding. The policy revision remained unchanged, but provider-environment refreshes produced new revision values on successive polls.

The tested compatibility triplet was mixed and is important context:

  • CLI/gateway: 0.0.117-dev.160+g9c41f057c
  • Supervisor image: metadata-only derivative of the dev.158 supervisor image; the only intended change was adding a private CA trust anchor
  • Sandbox runtime image: dev.158

If that combination is unsupported, sandbox creation should fail with an actionable compatibility error. It should not launch successfully and later fail through the boundary protocol.

I searched current issues and found related—but not exact—reports:

Impact / Why This Matters

The sandbox and its agent become unavailable even though the workload was healthy and the proxied request path had been working. The gateway terminates the workload after the control supervisor exits, so active sessions are lost and end-to-end automation cannot remain available.

The current diagnostics do not identify why the boundary became unavailable. The gateway records an empty Docker container wait error, while the earlier repeated endpoint-status failures are logged only as transient. This makes it difficult to distinguish a version-contract violation, boundary protocol failure, or recoverable transport interruption.

No automatic recovery occurred. Recreating the sandbox loses process/session state and prevents this configuration from being qualified for long-lived use.

Acceptance Criteria

  • A supported gateway/supervisor/runtime triplet remains healthy under normal allowed proxied traffic and periodic provider-environment refresh.
  • A transient Sandbox Protocol connection loss that successfully reconnects does not subsequently terminate the proxy accept loop or workload.
  • An unsupported gateway/supervisor/runtime protocol combination is rejected before workload launch with the incompatible components and revisions identified.
  • Repeated ReportEndpointStatus failures expose an actionable reason and do not remain an indefinitely ignored precursor to supervisor failure.
  • If the control supervisor must terminate, the sandbox condition retains the concrete boundary/supervisor error instead of an empty Docker container wait error.
  • A regression test covers boundary loss/recovery while the network-mediation proxy and provider-environment polling are active.

Reproduction Steps

The failure has been observed on a clean host but has not yet been reduced to a deterministic fault-injection trigger.

  1. Install OpenShell CLI/gateway 0.0.117-dev.160+g9c41f057c with the Docker driver.
  2. Select the dev.158 supervisor and sandbox-runtime images. Use a supervisor derivative that changes only the CA trust store.
  3. Create a sandbox using the split supervisor/workload boundary, automatic TLS inspection, Landlock best_effort, and two provider-projected environment values. No OpenShell content middleware is enabled.
  4. Start a long-running workload and make a normal policy-allowed proxied HTTPS request through host.openshell.internal.
  5. Leave the sandbox running. In the observed run, endpoint-status reports repeatedly returned gRPC status 9, followed after approximately 15 minutes by the boundary/proxy failure shown below.
  6. Inspect the sandbox: it transitions to Error with Ready=False, reason ControlSupervisorExited.

Environment

  • OpenShell CLI/gateway: 0.0.117-dev.160+g9c41f057c
  • Supervisor: dev.158 image, locally derived only to add a CA trust anchor
  • Sandbox runtime: dev.158 image
  • Compute driver: Docker
  • Docker: 29.8.0
  • OS: Ubuntu 24.04, x86_64
  • Kernel: 7.0.0-1012-aws
  • Policy/runtime: RFC 0012 split supervisor/workload, automatic TLS inspection, Landlock best_effort
  • Integration: APF-projected provider profiles with AgentGateway/ATGW; no OpenShell content middleware
  • Workload container after failure: OOMKilled=false, exit code 0 after the Docker driver stopped it
  • Latest tested build: dev.160. This RFC 0012 development layout was not destructively replaced with the older stable release solely for comparison.

Logs

# Repeated before the fatal event; config/provider fetches continued to succeed.
20:34:20.122Z ReportEndpointStatus http.response.status_code=200 rpc.grpc.status_code=9 otel.status_code="ERROR"
20:34:27.115Z GetSandboxConfig http.response.status_code=200
20:34:27.121Z GetSandboxProviderEnvironment completed successfully provider_count=2 env_count=2
20:34:27.124Z ReportEndpointStatus http.response.status_code=200 rpc.grpc.status_code=9 otel.status_code="ERROR"

# Fatal sequence.
20:34:35.799Z OCSF NET:FAIL [HIGH] [msg:Network-mediation source failed; proxy accept loop exiting: boundary unavailable: read boundary control response header: ...]
20:34:35.801Z WARN boundary_server::linux: Sandbox Protocol connection lost; workload frozen pending authenticated recovery
20:34:35.801Z OCSF FINDING:CREATE [MED] "Sandbox Supervisor Connection Lost" [type:sandbox-supervisor-connection-lost confidence:high]
20:34:35.805Z INFO boundary_server::linux: Sandbox Protocol connection recovered; workload resumed
20:34:45.858Z WARN boundary_server::linux: Boundary gRPC exchange failed
20:34:45.885Z INFO supervisor_session: supervisor session: stream closed by supervisor
20:34:45.885Z INFO supervisor_session: supervisor session: ended
20:34:45.886Z PushSandboxLogs http.response.status_code=200 rpc.grpc.status_code=13 latency_ms=909920
20:34:47.016Z WARN openshell_driver_docker: Failed to wait for Docker supervisor container error="Docker container wait error:"
20:34:47.032Z INFO compute: Sandbox phase changed old_phase=Provisioning new_phase=Error
20:34:47.032Z WARN compute: Sandbox failed to become ready reason=ControlSupervisorExited
20:34:47.046Z ERROR network_broker: sandbox network broker listener failed
20:34:47.109Z INFO openshell_driver_docker: Stopped Docker sandbox after control supervisor failure

# Workload container state after the driver stopped it.
state=exited exit=0 oom=false

All hostnames, sandbox/container identifiers, image digests, private endpoints, credentials, and enterprise payloads have been removed. Full sanitized logs can be provided if needed.

Activity

  1. andre-motta commented on Sep 22, 2026

    @andre-motta

    We hit the same failure path (network-mediation proxy exits with boundary unavailable, then the supervisor dies) with the Podman driver, and it reproduces deterministically without an agent.

    Version: gateway, CLI, supervisor and sandbox runtime built from v0.1.0-pre.5 (484f076). Rootless Podman inside a privileged container, x86_64.

    Reproducer

    1. Create a sandbox from any image that has uv and python3, with a base policy that allows uv to reach PyPI:

      openshell policy update --wait \
        --binary /usr/local/bin/uv \
        --add-endpoint pypi.org:443:read-only \
        --add-endpoint files.pythonhosted.org:443:read-only \
        <sandbox>
    2. Install a package with a large dependency tree:

      openshell sandbox exec --name <sandbox> --no-tty -- \
        sh -c 'uv venv -q /tmp/v && uv pip install -q -p /tmp/v mlflow'

    The exec fails partway through the install, and the supervisor container exits with code 1. It failed on the first attempt every time we tried.

    Error:   × code: 'The service is currently unavailable', message: "exec relay closed
      │ before the command reported an exit status"
    

    What the logs show

    Supervisor log, immediately before exit:

    OCSF CONFIG:PUBLISHED [INFO] Policy DNS mapped files.pythonhosted.org resolved=151.101.0.223,... synthetic=198.18.0.3 ports=443 mapping_id=...
    ... (dozens of ALLOWED GET http://files.pythonhosted.org:443/packages/... over ~30s)
    WARN openshell_supervisor_network::proxy: Denied staged transparent connection
    OCSF NET:OPEN [MED] DENIED /usr/local/bin/uv(0) -> 198.18.0.3:443 [reason:transparent_tcp_mapping_denied]
    OCSF NET:FAIL [HIGH] [msg:Network-mediation source failed; proxy accept loop exiting: boundary unavailable: read boundary control response: read o...]
    Error:   × control-mode proxy accept loop exited unexpectedly
    

    Workload container at the same moment:

    WARN openshell_sandbox::boundary_server::linux: Sandbox Protocol connection lost; workload frozen pending authenticated recovery
    WARN openshell_sandbox::boundary_server::linux: Boundary gRPC exchange failed   (x79)
    

    What we think is happening

    • uv resolves files.pythonhosted.org once, gets the synthetic IP, and keeps opening new connections to it for the rest of the install.
    • About 30 s after the lookup the policy-DNS mapping expires. The next connection to the cached synthetic IP is denied with transparent_tcp_mapping_denied, even though the host is allowed.
    • While other relayed connections are still open, the boundary control read fails right after that denial. The supervisor then treats Network-mediation source failed as fatal and exits (retain_remote_access_plane / the proxy_exited arm in openshell-supervisor/src/lib.rs), which takes down the workload.

    Each of these alone did not crash the supervisor in our testing:

    • a denied port on an allowed host
    • a single connect to an expired synthetic IP after sleeping 40 s
    • 40 parallel connects to an expired synthetic IP

    The crash needs the stale-mapping denial to happen while the install has live relays open.

    Notes

  2. politerealism commented on Sep 23, 2026

    @politerealism
    Contributor

    I'd like to work on this. @andre-motta's reproducer and root-cause hypothesis make this much more tractable than the original report.

    Before committing to a design, I want to run the reproducer against current main first — PR #3532 ("reclaim socket descriptors before exhaustion") merged yesterday and hasn't been tested against this specific failure yet, so it's worth confirming whether it already changes the outcome before designing around the stale-synthetic-IP-mapping theory.

    If the failure still reproduces on main, my read of the two directions suggested above (keep the policy-DNS mapping alive while a client may still be using it / re-resolve on connect, and don't let one failed boundary control read end the whole proxy accept loop) is that the second one needs care — the proxy exiting on a network-mediation failure is very likely an intentional fail-closed behavior, so any change there needs to distinguish "safe to recover from" versus "must still fail closed" rather than just relaxing the check. Will report back with findings before proposing a specific fix.

  3. added
    state:acceptedA maintainer decided OpenShell should pursue this issue
    and removed
    state:triage-neededOpened without agent diagnostics and needs triage
    on Sep 23, 2026
  4. added this to the OpenShell 0.1.0 milestone on Sep 23, 2026
  5. politerealism commented on Sep 23, 2026

    @politerealism
    Contributor

    Confirmed: PR #3532 does not fix this. Reproduced @andre-motta's scenario (Podman, rootless, uv pip install mlflow with a policy allowing pypi.org/files.pythonhosted.org) on current main (commit includes #3532), and hit the identical failure:

    Error:   × control-mode proxy accept loop exited unexpectedly
    

    Same end-to-end signature as the original report and @andre-motta's comment:

    • Workload boundary log spam: Boundary gRPC exchange failed
    • Gateway: exec relay closed before the command reported an exit status
    • Sandbox phase: Provisioning → Error

    One gap on my end: my supervisor log was too sparse (only 3 lines total) to confirm the exact trigger @andre-motta identified (the stale policy-DNS mapping denial via transparent_tcp_mapping_denied) — I only captured the end-state crash, not the per-request OCSF detail leading up to it. Planning to bump log verbosity and rerun to confirm that specific trigger before proposing a fix design, since the two candidate directions (keep the DNS mapping alive while in use vs. not letting one failed boundary read kill the whole proxy loop) need that confirmed rather than assumed.

    Ruling out #3532 as a fix means the fail-closed-behavior question is still the real open design work here.

  6. moved this from Todo to In progress in OpenShell Roadmapon Sep 23, 2026
  7. politerealism commented on Sep 23, 2026

    @politerealism
    Contributor

    @drew's PR #3642 nails this — confirmed by our own independent hard evidence, not just taking the PR's word for it.

    We reproduced the same crash with RUST_LOG=info,openshell_sandbox_backend=trace,h2=trace,tonic=trace on the supervisor and found:

    • Zero transparent_tcp_mapping_denied/_expired events in the entire crash trace. The stale-DNS-mapping theory (ours and the original report's) doesn't hold up — matches PR fix(sandbox): keep boundary connection live under stalled relays #3642's explicit finding: "The 30-second policy DNS mapping expiry is not the cause."
    • At the moment of failure, ~175 concurrent HTTP/2 streams (StreamId up to 349) all transitioned to Closed(Error(Io(BrokenPipe, ...))) within a 4ms window, starting from StreamId(1). That's the signature of the entire shared connection being torn down at once, not an independent per-stream network failure — consistent with fix(sandbox): keep boundary connection live under stalled relays #3642's root cause: recover_after_unavailable drops the whole channel on a single stalled-relay-triggered stream failure, then the reattach race (rejected due to the sandbox not yet observing the old connection's disconnect) turns a recoverable wedge into a fatal supervisor exit.

    Good outcome: the investigation converged on the same place from two independent directions (connection-level flow-control starvation + reconnect race, not policy/DNS). Filing a separate issue for the Podman /etc/resolv.conf gap noted in #3642's follow-ups, since we hit that exact same gap independently during our own reproduction attempts and it's explicitly unresolved by this PR.

  8. moved this from In progress to Done in OpenShell Roadmapon Sep 23, 2026
  9. added a commit that references this issue on Sep 24, 2026
    490055b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

state:acceptedA maintainer decided OpenShell should pursue this issue

Type

No type

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions