Repository navigation
Sandbox stuck in unrecoverable Error phase after supervisor session drops (SSH disconnect); stop/start refuse to act #3308
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 14, 2026 Correction: root cause identified, causality was reversed
We initially assumed the sandbox's supervisor session died because the client's SSH connection dropped first (a network blip). Reproducing this twice in separate sandboxes shows the opposite: the SSH disconnect is a symptom, not the cause.
In both runs, the exact same sequence occurs, timestamps ~1s apart:
- A policy-denied outbound connection attempt occurs (in our repro: OpenClaw's onboarding flow reaching
registry.npmjs.org:443, correctly denied by policy). The deny log shows a peer-binary resolution failure:NET:OPEN [MED] DENIED -> registry.npmjs.org:443 [engine:opa] [reason:failed to resolve peer binary: No ESTABLISHED TCP connection found for 10.200.0.2:46932 -> 10.200.0.1:3128 in /proc/53/net/tcp{,6}] - Under 1 second later, the supervisor session itself errors and ends:
[gateway] [WARN] relay stream: inbound errored [gateway] [WARN] supervisor session: stream error [gateway] [INFO] supervisor session: ended
That supervisor-session death is what the SSH terminal is riding on — which is what shows up locally as
client_loop: send disconnect: Broken pipe. So the causality is the reverse of what we originally assumed: it's not your SSH connection dying and killing the sandbox — it's a bug in how the sandbox's network-policy engine handles a peer-binary resolution failure during a DENY that's crashing the supervisor session (and therefore the sandbox itself).See the new issue (linked below) for the crash-on-deny root cause. This issue (unrecoverable Error phase, no stop/start recovery path) still stands on its own regardless of what triggers entry into Error phase, but the reproduction steps and "Expected" section describing it as caused by a client-side SSH drop should be disregarded.
- A policy-denied outbound connection attempt occurs (in our repro: OpenClaw's onboarding flow reaching
We hit the Error-phase recovery block on OpenShell
0.0.116after deleting a test sandbox Pod. The environment was Agent Sandbox1.0.2, K3sv1.34.11+k3s1, and NVIDIA/nemoclaw-community#166 atcf03dcdd38e6a6627a9946b6cc9f4d9125c892f9.- We deleted only the Pod of a working sandbox, leaving its native Sandbox CR intact.
- The replacement got a new UID and became Ready. The native Sandbox CR reported Ready/DependenciesReady.
- The access relay stayed unavailable. OpenShell reported
Error, and authenticated supervised exec refused to run. sandbox stopreturnedsandbox must be Ready to stop (current phase: Error). We didn't teststartfrom Error.
I read the correction to the original report. Our trigger was Pod deletion, not an SSH disconnect, and we haven't established a shared root cause with the reported policy-deny failure.
The recipe's
lifecycle.sandbox.desiredState=absent, thenpresentworkflow restored service and preserved the external PVC identities. That recreated the logical sandbox; automatic recovery did not work in this test.Should this go here or under #2616? Unlike its Pending/Provisioning example, our replacement was already Kubernetes Ready. We're looking for a supported recovery path that retains durable storage, not restoration of running processes or active sessions.
Reproduced the unrecoverable
Errorphase on 0.1.2 with the Docker driver, triggered by a gateway restart rather than an SSH drop:- Two
Readysandboxes on a local gateway. - Restart the gateway with a config under which supervisors can't reach it (on WSL 2 + Docker Desktop, removing the
grpc_endpointworkaround from bug(docker): sandboxes never become Ready on WSL 2 + Docker Desktop (supervisor dials 127.0.0.1) #3880 does this). - Both sandboxes go to
ErrorwithReady: False (StartFailed) - Failed to start sandbox during gateway startup: Docker supervisor exited before becoming ready ... failed to connect to OpenShell server. - Restore the working config and restart the gateway. New sandboxes reach
Ready, but the existing two stay inError:
$ openshell sandbox stop e2e-multi Error: ... sandbox must be Ready to stop (current phase: Error) $ openshell sandbox start e2e-multi Error: ... sandbox must be Stopped, Completed, or a failed main-process Error to start (current phase: Error)So on 0.1.2, a transient gateway-side startup failure (
StartFailed, not a main-process failure) still leaves no recovery path short of delete. RetryingStartFailedsandboxes once the gateway is healthy, or allowingstartfor them, would cover this case.- Two
Just noting we are running into this too. Any temporary issue on the k8s side and there's not a great recovery story there. No suggestions at this point - just wanted to chime in to help reinforce this is a legitimate problem that could use some attention soon.
- addedstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issuestate:needs-infoAssessment needs specific evidence or reproduction detailsAssessment needs specific evidence or reproduction detailsand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Oct 6, 2026 @stmcginnis can you elaborate on what type of issues you've seen on the k8s side that are causing this. also confirming it's on 0.1.x?
Hey @johntmyers - thanks for checking!
Current environment is v0.1.2 running on a k8s cluster using Kata for microVMs on metal instances.
The trigger for this is (or the main one I'm seeing at least) is losing a node under the sandbox. So a host gets replaced or drained, or maybe it crashes. Something causes an eviction, etc. The sandbox pods run with
restartPolicy: Never, so the pod endsFailedand is never replaced.The gateway then records the sandbox as
ErrorwithReady=False,reason=PodFailed, andSuspended=Truesoon after. The workspace PVC survives with the user's files intact, at least whatever was flushed to disk before going down.startrefuses becauseis_failed_main_process_resultonly acceptsMainProcessFailedwith an exit code, andstoprefuses because the phase isn'tReady. The only way out is to delete it, which loses the PVC.Aloowing
startforPodFailedandStartFailed, in addition to the existingMainProcessFailed, I think would give a path to get back running again.Can verify that I'm seeing this as well. Seems like anything that removes a pod results in this.
@stmcginnis @johntmyers could I ask y'all to look at #4231 which aims to fix this? I've sic'd gator on it as well.
- addedarea:gatewayGateway server and control-plane workGateway server and control-plane workos:linuxIssue affects Linux hostsIssue affects Linux hostsand removedstate:needs-infoAssessment needs specific evidence or reproduction detailsAssessment needs specific evidence or reproduction detailsstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issue
on Oct 8, 2026
Summary
When the SSH connection carrying a sandbox's supervisor session drops (e.g. client-side "Broken pipe"), the sandbox transitions to
Phase: Errorand there is no way to recover it short ofsandbox delete+sandbox create.sandbox stopandsandbox startboth refuse to act on anError-phase sandbox, so the workspace inside/sandboxis unrecoverable.Steps to reproduce
openshell sandbox create(image/policy runningopenclaw-startas main command)client_loop: send disconnect: Broken pipe)openshell sandbox get <name>→Phase: Erroropenshell sandbox stop <name>→Error: sandbox must be Ready to stop (current phase: Error)openshell sandbox start <name>→Error: sandbox must be Stopped to start (current phase: Error)Relevant logs
(log capture was truncated at time of report)
Expected
Either:
Errorphase should have a documented recovery path (e.g.sandbox stop --force/sandbox recover) that doesn't discard the workspace.Actual
The only path out of
Errorphase issandbox delete, which discards the sandbox's workspace with no export/backup step offered.openshell version: 0.0.116