Repository navigation
Failed peer-binary resolution during policy DENY crashes the supervisor session, tearing down the sandbox's SSH relay #3311
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 14, 2026 📋 triage-agent
Assessment: historical sandbox/session loss reported on 0.0.116; the proposed crash-on-identity-denial mechanism is not confirmed on current main. Keep open pending current reproduction evidence.
Reviewed
mainat5601d71b014545678c9fcb572dad7e6cc09fc448on October 6, 2026, and compared the relevant path with tagv0.0.116.- The quoted “No ESTABLISHED TCP connection” error is a process-identity lookup failure. It does not establish that OPA evaluated the destination allowlist and denied it.
- Both 0.0.116 and current main convert that lookup error into a
NetworkAction::Deny. The explicit CONNECT handler sends HTTP 403 and returns; connection-handler errors are logged in an individual task. I did not find the claimed direct error propagation from this denial into supervisor-session teardown. - Current main additionally handles backend-supplied identity failures as per-open
IdentityUnavailabledenials. The sandbox broker returns a connection errno to the workload. Source coverage includes identity-unavailable denial and denied-connect behavior, but these tests do not establish that the original OpenClaw/session-loss scenario is fixed. - Relevant changes since the report include the runtime/boundary refactor in feat(isolation): implement the RFC 0012 sandbox architecture #2942, authenticated TCP mediation recovery in fix(sandbox-backend): recover TCP mediation after boundary disconnects #3403, startup session retries in fix(supervisor): keep session retries during startup #3765, and SSH relay replay/recovery in fix(runtime): recover SSH relays and bound startup diagnostics #4011. These improve resilience; none establishes that this exact reported failure is resolved.
Source references: identity lookup failure handling, CONNECT denial response, per-open identity denial, workload connection error.
To establish current relevance, please provide a reproduction on current main or a recent release with the exact gateway/runtime versions, driver, workload image/tag, policy and command, complete gateway and supervisor logs around the incident, and container/Pod termination reason and exit code. In particular, distinguish supervisor/runtime termination from the workload exiting after its network request fails.
Validation here was source/history inspection; I did not run a fresh OpenClaw reproduction or execute the cited tests. The two original runs support investigating session loss, but their timestamp correlation alone does not establish the proposed root cause. #3308 remains the separate recovery-path issue.
Next human decision: with current reproduction evidence, validate and scope the session-loss defect; otherwise decide whether the historical report can be closed. I am not marking this fixed from code inspection alone.
- addedstate:needs-infoAssessment needs specific evidence or reproduction detailsAssessment needs specific evidence or reproduction detailsand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Oct 6, 2026 @papillonlibre can you please determine if this is still relevant on 0.1.x and report back
Summary
A single correctly denied outbound request — not a bypass attempt, not a policy misconfiguration, just an app trying to reach a host that isn't allowlisted — can kill the whole sandbox. That's a significant reliability/security-usability problem: the failure mode of "policy correctly blocks something" should be "request fails," not "sandbox dies."
Steps to reproduce
openshell sandbox createwith a policy that does not allowlist a host the sandbox's main command will try to reach (in our repro:openclaw-start, whose onboarding flow reachesregistry.npmjs.org:443, not present in the default policy'snvidia/nvidia_web/github/github_rest_api/gitlab/claude_codenetwork_policies).openshell logs <name>.client_loop: send disconnect: Broken pipe).openshell sandbox get <name>→Phase: Error, no recovery available.Logs (two independent reproductions)
Run 1 (sandbox 42449366-d949-4285-b787-aa7415570ccf):
Run 2 (sandbox 9e489e30-42f3-411a-983d-...):
Expected
A policy DENY — including one where peer-binary resolution fails — should result in the connection attempt failing cleanly from the sandboxed process's point of view. It should never crash the supervisor session or bring down the sandbox.
Suspected area
The peer-binary resolution path in the OPA-backed network policy engine (
openshell_supervisor_network::opa), specifically the branch when no matching entry is found in/proc/<pid>/net/tcp{,6}for the connecting socket — this looks like a race (the connecting process's socket may have already closed/reused by the time policy inspects/proc) that isn't handled gracefully and instead propagates into the supervisor session, killing it.openshell version: 0.0.116
Related: #3308