Repository navigation
bug: SSH server accept loop exits permanently on transient errors (same class as #2337) #2372
Copy link
Copy link
Closed
Bug
Copy link
Labels
area:supervisorProxy and routing-path workProxy and routing-path workstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.topic:observabilityLogging, metrics, and observability workLogging, metrics, and observability work
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Jul 20, 2026 📋 triage-agent
Triage Assessment
Classification: bug-confirmed
Summary
A transient SSH accept error can terminate the accept task while the sandbox continues to appear ready.
Investigation
The SSH loop still propagates accept failures with ?, so the task can exit without repairing readiness. Related PRs #2369 and #2370 address proxy behavior, not this SSH loop.
Recommendation
Add bounded retry for transient resource exhaustion and supervise unexpected task exit; do not blindly retry permanent listener failures.
- addedarea:supervisorProxy and routing-path workProxy and routing-path worktopic:observabilityLogging, metrics, and observability workLogging, metrics, and observability workand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Jul 23, 2026 This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 9, 2026 Working on this now.
- added 2 commits that reference this issue
on Aug 11, 2026 - added 2 commits that reference this issue
on Aug 25, 2026 - added a commit that references this issue
on Aug 25, 2026
Metadata
Metadata
Assignees
Labels
area:supervisorProxy and routing-path workProxy and routing-path workstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.topic:observabilityLogging, metrics, and observability workLogging, metrics, and observability work
Agent Diagnostic
Identified during review of PRs #2369 and #2370 (fix for #2337). Three independent review agents flagged the same vulnerability in the SSH server accept loop.
Finding:
crates/openshell-supervisor-process/src/ssh.rs:141uses?onlistener.accept().await, so any transient error (EMFILE, ENFILE, ECONNABORTED) propagates out ofrun_ssh_server()and permanently kills the SSH accept loop. The spawned task incrates/openshell-supervisor-process/src/run.rs:238is fire-and-forget — theJoinHandleis dropped, so the sandbox has no way to detect that the SSH server has exited.This is the same class of bug as #2337, but in the exec relay path instead of the proxy path.
Description
What happens: When the SSH server's TCP listener hits a transient error like EMFILE (too many open file descriptors), the accept loop exits permanently. The sandbox continues to report
Readyeven though all subsequent exec relays will fail. The OCSF event is emitted atCriticalseverity, but no recovery or sandbox termination occurs.What should happen: The SSH server should either:
Why this matters: The proxy and SSH server share the supervisor's FD table. FD exhaustion triggered by proxy connections (the scenario in #2337) can simultaneously kill the SSH server. PR #2369 keeps the proxy alive through backoff, but during the backoff period the SSH accept can still fail and exit permanently. The proxy recovers, but exec relays are dead.
This means the invariant from #2337 — "Ready must imply the sandbox's required policy proxy and exec relay are still operational" — is only enforced for the proxy path, not the SSH/exec relay path.
Reproduction Steps
listener.accept().awaitreturns EMFILE?operator propagates the error, exitingrun_ssh_server()openshell execcommands fail because the SSH server is goneEnvironment
crates/openshell-supervisor-process/src/ssh.rs:141,crates/openshell-supervisor-process/src/run.rs:238-262Suggested Fix
Apply the same two-layer defense pattern from PRs #2369 and #2370:
Retry with backoff (
ssh.rs): Replacelistener.accept().await.into_diagnostic()?with a sleep-and-continue pattern matchingproxy.rs. Reuse or mirroris_fd_exhaustion_error()andaccept_backoff().Exit notification (
ssh.rs+lib.rs): Add a oneshot drop-guard channel (matching the_proxy_exit_guardpattern from PR fix(sandbox): terminate sandbox when proxy accept loop exits unexpectedly #2370). Store theJoinHandleor exit receiver and race it in the sandbox'stokio::select!wait paths.Related