Repository navigation
bug(supervisor-network): boundary reconnect is followed by proxy exit and ControlSupervisorExited #3396
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 16, 2026 - added a commit that references this issue
on Sep 17, 2026 - added a commit that references this issue
on Sep 22, 2026 We hit the same failure path (network-mediation proxy exits with
boundary unavailable, then the supervisor dies) with the Podman driver, and it reproduces deterministically without an agent.Version: gateway, CLI, supervisor and sandbox runtime built from
v0.1.0-pre.5(484f076). Rootless Podman inside a privileged container, x86_64.Reproducer
-
Create a sandbox from any image that has
uvandpython3, with a base policy that allowsuvto reach PyPI:openshell policy update --wait \ --binary /usr/local/bin/uv \ --add-endpoint pypi.org:443:read-only \ --add-endpoint files.pythonhosted.org:443:read-only \ <sandbox>
-
Install a package with a large dependency tree:
openshell sandbox exec --name <sandbox> --no-tty -- \ sh -c 'uv venv -q /tmp/v && uv pip install -q -p /tmp/v mlflow'
The exec fails partway through the install, and the supervisor container exits with code 1. It failed on the first attempt every time we tried.
Error: × code: 'The service is currently unavailable', message: "exec relay closed │ before the command reported an exit status"What the logs show
Supervisor log, immediately before exit:
OCSF CONFIG:PUBLISHED [INFO] Policy DNS mapped files.pythonhosted.org resolved=151.101.0.223,... synthetic=198.18.0.3 ports=443 mapping_id=... ... (dozens of ALLOWED GET http://files.pythonhosted.org:443/packages/... over ~30s) WARN openshell_supervisor_network::proxy: Denied staged transparent connection OCSF NET:OPEN [MED] DENIED /usr/local/bin/uv(0) -> 198.18.0.3:443 [reason:transparent_tcp_mapping_denied] OCSF NET:FAIL [HIGH] [msg:Network-mediation source failed; proxy accept loop exiting: boundary unavailable: read boundary control response: read o...] Error: × control-mode proxy accept loop exited unexpectedlyWorkload container at the same moment:
WARN openshell_sandbox::boundary_server::linux: Sandbox Protocol connection lost; workload frozen pending authenticated recovery WARN openshell_sandbox::boundary_server::linux: Boundary gRPC exchange failed (x79)What we think is happening
uvresolvesfiles.pythonhosted.orgonce, gets the synthetic IP, and keeps opening new connections to it for the rest of the install.- About 30 s after the lookup the policy-DNS mapping expires. The next connection to the cached synthetic IP is denied with
transparent_tcp_mapping_denied, even though the host is allowed. - While other relayed connections are still open, the boundary control read fails right after that denial. The supervisor then treats
Network-mediation source failedas fatal and exits (retain_remote_access_plane/ theproxy_exitedarm inopenshell-supervisor/src/lib.rs), which takes down the workload.
Each of these alone did not crash the supervisor in our testing:
- a denied port on an allowed host
- a single connect to an expired synthetic IP after sleeping 40 s
- 40 parallel connects to an expired synthetic IP
The crash needs the stale-mapping denial to happen while the install has live relays open.
Notes
- fix(sandbox-backend): recover TCP mediation after boundary disconnects #3403 and fix(sandbox-backend): confirm renewed boundary credentials within an epoch #3411 are included in
v0.1.0-pre.5and do not prevent this. - fix(sandbox): reclaim socket descriptors before exhaustion #3532 (reclaim socket descriptors before exhaustion) looks related, since it describes socket-heavy workloads taking down the control connection. It landed after
v0.1.0-pre.5, and we have not tested it yet. - Two things would each make this survivable: keep a policy-DNS mapping alive while a client may still be using it (or re-resolve the synthetic IP on connect), and don't let one failed boundary control read end the whole proxy accept loop.
-
- added a commit that references this issue
on Sep 22, 2026 I'd like to work on this. @andre-motta's reproducer and root-cause hypothesis make this much more tractable than the original report.
Before committing to a design, I want to run the reproducer against current
mainfirst — PR #3532 ("reclaim socket descriptors before exhaustion") merged yesterday and hasn't been tested against this specific failure yet, so it's worth confirming whether it already changes the outcome before designing around the stale-synthetic-IP-mapping theory.If the failure still reproduces on
main, my read of the two directions suggested above (keep the policy-DNS mapping alive while a client may still be using it / re-resolve on connect, and don't let one failed boundary control read end the whole proxy accept loop) is that the second one needs care — the proxy exiting on a network-mediation failure is very likely an intentional fail-closed behavior, so any change there needs to distinguish "safe to recover from" versus "must still fail closed" rather than just relaxing the check. Will report back with findings before proposing a specific fix.- addedstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issueand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 23, 2026 Confirmed: PR #3532 does not fix this. Reproduced @andre-motta's scenario (Podman, rootless,
uv pip install mlflowwith a policy allowingpypi.org/files.pythonhosted.org) on currentmain(commit includes #3532), and hit the identical failure:Error: × control-mode proxy accept loop exited unexpectedlySame end-to-end signature as the original report and @andre-motta's comment:
- Workload boundary log spam:
Boundary gRPC exchange failed - Gateway:
exec relay closed before the command reported an exit status - Sandbox phase:
Provisioning→Error
One gap on my end: my supervisor log was too sparse (only 3 lines total) to confirm the exact trigger @andre-motta identified (the stale policy-DNS mapping denial via
transparent_tcp_mapping_denied) — I only captured the end-state crash, not the per-request OCSF detail leading up to it. Planning to bump log verbosity and rerun to confirm that specific trigger before proposing a fix design, since the two candidate directions (keep the DNS mapping alive while in use vs. not letting one failed boundary read kill the whole proxy loop) need that confirmed rather than assumed.Ruling out #3532 as a fix means the fail-closed-behavior question is still the real open design work here.
- Workload boundary log spam:
@drew's PR #3642 nails this — confirmed by our own independent hard evidence, not just taking the PR's word for it.
We reproduced the same crash with
RUST_LOG=info,openshell_sandbox_backend=trace,h2=trace,tonic=traceon the supervisor and found:- Zero
transparent_tcp_mapping_denied/_expiredevents in the entire crash trace. The stale-DNS-mapping theory (ours and the original report's) doesn't hold up — matches PR fix(sandbox): keep boundary connection live under stalled relays #3642's explicit finding: "The 30-second policy DNS mapping expiry is not the cause." - At the moment of failure, ~175 concurrent HTTP/2 streams (StreamId up to 349) all transitioned to
Closed(Error(Io(BrokenPipe, ...)))within a 4ms window, starting from StreamId(1). That's the signature of the entire shared connection being torn down at once, not an independent per-stream network failure — consistent with fix(sandbox): keep boundary connection live under stalled relays #3642's root cause:recover_after_unavailabledrops the whole channel on a single stalled-relay-triggered stream failure, then the reattach race (rejected due to the sandbox not yet observing the old connection's disconnect) turns a recoverable wedge into a fatal supervisor exit.
Good outcome: the investigation converged on the same place from two independent directions (connection-level flow-control starvation + reconnect race, not policy/DNS). Filing a separate issue for the Podman
/etc/resolv.confgap noted in #3642's follow-ups, since we hit that exact same gap independently during our own reproduction attempts and it's explicitly unresolved by this PR.- Zero
- added a commit that references this issue
on Sep 24, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
User Story
As an operator running a long-lived agent sandbox through OpenShell's RFC 0012 split supervisor/workload architecture, I need a transient Sandbox Protocol boundary failure to recover cleanly—or be rejected as an incompatible deployment before launch—so that an otherwise healthy workload is not terminated and left in
Error.Problem Statement
A Docker-backed sandbox ran successfully for approximately 15 minutes, including allowed proxied HTTPS traffic and successful in-place gateway/Sandbox Protocol credential renewal. The network-mediation proxy then exited with
boundary unavailable.The workload boundary detected the supervisor connection loss, froze the workload, and reported recovery about 4 ms later. Ten seconds after that apparent recovery, the boundary gRPC exchange failed. The supervisor stream ended, the gateway reported an empty Docker wait error, and the sandbox transitioned to
Errorwith reasonControlSupervisorExited. The Docker driver then stopped the workload; the workload container itself was not OOM-killed and exited 0.Repeated
ReportEndpointStatuscalls were already returning gRPC status 9 before the fatal event, whileGetSandboxConfigandGetSandboxProviderEnvironmentcontinued succeeding. The policy revision remained unchanged, but provider-environment refreshes produced new revision values on successive polls.The tested compatibility triplet was mixed and is important context:
0.0.117-dev.160+g9c41f057cIf that combination is unsupported, sandbox creation should fail with an actionable compatibility error. It should not launch successfully and later fail through the boundary protocol.
I searched current issues and found related—but not exact—reports:
Errorstate after a supervisor session ends.Readyrather than enteringError.EMFILE/ENFILEerror was observed here.Impact / Why This Matters
The sandbox and its agent become unavailable even though the workload was healthy and the proxied request path had been working. The gateway terminates the workload after the control supervisor exits, so active sessions are lost and end-to-end automation cannot remain available.
The current diagnostics do not identify why the boundary became unavailable. The gateway records an empty
Docker container wait error, while the earlier repeated endpoint-status failures are logged only as transient. This makes it difficult to distinguish a version-contract violation, boundary protocol failure, or recoverable transport interruption.No automatic recovery occurred. Recreating the sandbox loses process/session state and prevents this configuration from being qualified for long-lived use.
Acceptance Criteria
ReportEndpointStatusfailures expose an actionable reason and do not remain an indefinitely ignored precursor to supervisor failure.Docker container wait error.Reproduction Steps
The failure has been observed on a clean host but has not yet been reduced to a deterministic fault-injection trigger.
0.0.117-dev.160+g9c41f057cwith the Docker driver.best_effort, and two provider-projected environment values. No OpenShell content middleware is enabled.host.openshell.internal.ErrorwithReady=False, reasonControlSupervisorExited.Environment
0.0.117-dev.160+g9c41f057c29.8.024.04, x86_647.0.0-1012-awsbest_effortOOMKilled=false, exit code0after the Docker driver stopped itLogs
All hostnames, sandbox/container identifiers, image digests, private endpoints, credentials, and enterprise payloads have been removed. Full sanitized logs can be provided if needed.