Repository navigation
feat: support ephemeral job-style sandboxes #2879
Description
Activity
The agent called this a feature, but it's actually a regression.
This workflow used to work:
> openshell sandbox create -- echo hi Created sandbox: prudent-werewolf hiBut now it returns an error:
> openshell sandbox create -- echo hi Created sandbox: pious-whitebait ✗ Error: MainProcessExited: Canonical main process exited Error: × sandbox entered error phase while provisioning: MainProcessExited: Canonical main process exited- addedarea:cliCLI-related workCLI-related workarea:sandboxSandbox runtime and isolation workSandbox runtime and isolation workarea:gatewayGateway server and control-plane workGateway server and control-plane workarea:supervisorProxy and routing-path workProxy and routing-path worktopic:compatibilityCompatibility-related workCompatibility-related worktest:e2eRequires end-to-end coverageRequires end-to-end coverage
on Aug 21, 2026 📋 triage-agent
Triage Assessment
Classification: validated-bug (capability regression; remediation may be feature-shaped)
Summary
The finite-command workflow is a confirmed regression introduced by PR #2726 (
ef296806), which made trailing create arguments the persistent canonical main process and intentionally made every canonical-process exit terminal, including exit code 0. The user requirement is the observable job contract: run one isolated finite command, return its output and exit status, and clean up reliably. Preserving the formersandbox create -- ...contract is preferred but not required; an explicit job flag, command, or API mode is acceptable.Investigation
- Before
ef296806, trailing create arguments ran through the independentsandbox execpath after provisioning. The E2E contract required bothsandbox create -- echo OKandsandbox create --no-keep -- echo OKto succeed; the latter then deleted the sandbox. - PR feat(sandbox): add canonical main process #2726 moved the trailing command into
SandboxSpec.command. Current gateway code records its exit code, setsMainProcessExited, and transitions the sandbox to terminalErrorregardless of whether the exit code is 0. - Current E2E coverage explicitly requires
sandbox create --no-keep -- echo OKto fail and then clean up. This directly corroborates the reported result. The CLI help still describes--no-keepas deleting the sandbox after the initial command or shell exits. sandbox execalready returns command output and exit status, but create → exec → delete is client-coordinated and non-atomic. There is no first-class CLI or API job lifecycle.- The requested semantics are technically coherent. Existing command transport, exit reporting, output streaming/replay, and cleanup primitives provide most building blocks. The main contract choices are the job selector, exact CLI exit-code propagation, terminal API representation/result retention, output retention limits, and cleanup/cancellation timing.
- No duplicate was found. feat: define a canonical sandbox main process and reconnectable session #2710 is the causal persistent-main design; feat: support custom sandbox entrypoint command #848 and sandbox doesn't run the initial custom command after system restart #1805 are related but cover different behavior.
- The report does not identify an exact OpenShell version. v0.0.110 predates the causal merge; the current tree and local v0.0.111 tag contain it. Missing version data limits released-user scope, not technical validity.
Impact Signals
- Affected users/scope: Automation using finite commands during sandbox creation; shared gateway/supervisor semantics affect Docker, Podman, Kubernetes, and VM contracts, though cross-driver conformance still needs E2E verification.
- Regression: Yes; the prior E2E success contract changed deliberately in
ef296806. - Workaround: Available (
sandbox create→sandbox exec→sandbox delete) but non-atomic, requires client lifecycle coordination, and can leak resources when interrupted. - Evidence quality: High for current behavior and causality (report reproduction, current code, pre/post E2E history); exact shipped-version scope is unknown.
Human Decision Required
Decide whether OpenShell should address this issue. If yes, apply
state:accepted, associate it with a roadmap item, or do both, and decide whether the work remains human-owned. Either action records acceptance; roadmap placement additionally records sequencing. To queue investigation or planning for an unattended agent, also applyagent:plan-requested. You can instead directly ask an agent to usecreate-spikeorbuild-from-issueon this issue. If no, close it as not planned and record the rationale.- Before
🏗️ build-plan
Implementation Plan
Issue type:
fix
Complexity: High
Confidence: Medium-high — lifecycle mapping is clear; short-command output needs a bounded terminal-drain handshake.Summary
Restore the existing
sandbox create -- <command>finite-command contract without adding a job RPC or selector. The canonical main process remains lifecycle authority: exit 0 becomesCompleted, nonzero becomesStopped/MainProcessFailed, andErrorremains reserved for infrastructure failures. Foreground create streams the existing main session and propagates its exact exit status;--no-keepdeletes after output/result drain.Scope
proto/openshell.protoand SDKs: appendSANDBOX_PHASE_COMPLETED = 9; update conversions and rendering without renumbering existing values.crates/openshell-server: persist zero/nonzero main results with distinct phases and conditions; make Completed sticky, watch-terminal, restartable, and safe against stale exit and snapshot races.crates/openshell-supervisor-processand relay authorization: keep short-process output attachable through a bounded terminal-drain window and preserve exact stdout, stderr, and TTY behavior.crates/openshell-cli: attach explicit commands regardless of terminal detection unless--detach; stream output, return the main exit code, and delete--no-keepsandboxes after drain.- E2E fixtures: keep fix(sandbox): stabilize canonical main process tests #2854's VM argv encoding; retain durable commands only where tests truly require a live sandbox; restore finite-command coverage elsewhere.
- Architecture, published sandbox lifecycle docs, and the OpenShell CLI skill: document Completed and foreground/ephemeral behavior.
Implementation Steps
- Add and propagate the additive Completed enum value through Rust, Go, TypeScript, Python-facing generated bindings, CLI, and TUI.
- Change gateway exit mapping and lifecycle reconciliation: zero → Completed/MainProcessCompleted; nonzero → Stopped/MainProcessFailed; preserve exact normalized exit code and stale-instance fencing.
- Make early completion attachable: establish and acknowledge the supervisor relay, preserve the bounded MainSession replay buffer through terminal drain, and narrowly allow matching terminal-main attachments.
- Update foreground create, exact exit propagation, and
--no-keeppost-drain cleanup using existing APIs. - Selectively unwind fix(sandbox): stabilize canonical main process tests #2854/test(e2e): align detached sandbox assertions #2856 durable test workarounds and add zero, nonzero, fast-exit, detach, restart, and cleanup coverage.
- Update architecture, docs, and CLI skill; run consistency checks and verification.
Test Plan
- Unit: phase conversions and rendering; zero/nonzero exit mapping; Completed lifecycle precedence, watch termination, restart and stale-exit behavior; terminal-drain notification and timeout.
- Integration: foreground create streams and returns 0 or exact nonzero; explicit detach remains Ready;
--no-keepdeletes after success and failure. - E2E: lifecycle paths across supported drivers; Kubernetes namespace and OIDC fixtures; MCP reusable sandbox remains explicitly detached.
Risks & Open Questions
- Old clients may render additive phase value 9 as Unknown; coordinated SDK updates and documentation are required.
- Main output replay remains bounded to the existing 1 MiB ring before attachment.
- Completed must not be overwritten by late supervisor disconnects or driver snapshots, and cleanup must not race output drain.
- Detached short jobs require a bounded drain timeout so compute terminates promptly.
Documentation Impact
Update
architecture/sandbox.md,architecture/gateway.md,architecture/compute-runtimes.md,docs/sandboxes/manage-sandboxes.mdx, relevant compute-driver guidance, and.agents/skills/openshell-cli/. No gateway TOML or Helm configuration documentation change is required. LSM behavior is unchanged because process identity and enforcement boundaries remain unchanged.
Revision 1 — Completed phase and existing-API foreground command contract
Implemented in #2884.
The final contract is:
- canonical main exit 0 ->
CompletedwithMainProcessCompletedand exit code 0 - canonical main nonzero/signal exit ->
StoppedwithMainProcessFailedand the exact normalized exit code - infrastructure failures remain
Error - an explicit trailing command stays foreground and streams output unless
--detachis set --no-keepdrains output/result before deletion, with a gateway-side ephemeral cleanup fallback- retained
Completedand failedStoppedsandboxes can be started again
The supervisor continues to own the canonical process. It now registers and retains the main-output relay long enough for fast commands to attach and drain, rather than suppressing the relay when the command exits before SSH is ready.
Validation includes the full server unit suite, pre-commit, SDK tests, and Docker lifecycle E2E covering successful and failed commands, output streaming, retained terminal state, reconnect replay, and ephemeral cleanup.
- canonical main exit 0 ->
User Story
As an automation user, I want to create an ephemeral sandbox that runs one finite command, so that I can execute isolated jobs and determine their outcome from the command's exit status without managing a long-lived sandbox.
Problem Statement
OpenShell's canonical main-process lifecycle treats every main-process exit as a terminal sandbox error, including a successful exit with code 0. This is appropriate for persistent, deployment-style sandboxes whose canonical process is expected to remain running, but it prevents finite commands from completing successfully as part of sandbox creation.
OpenShell needs an explicit job-style lifecycle for finite workloads, distinct from the existing persistent lifecycle established by #2710.
Impact / Why This Matters
One-shot commands such as build steps, tests, batch processing, and automation tasks cannot use sandbox creation as a successful finite operation. A command that completes normally is reported as a runtime failure, so callers cannot distinguish successful job completion from provisioning or runtime failure.
The current workaround is to create a retained sandbox, run the command with
sandbox exec, capture its result, and then delete the sandbox manually. That workflow is not atomic, requires extra lifecycle coordination, can leak resources when automation is interrupted, and does not provide a single job result.Proposed Design
Allow users to explicitly create an ephemeral, job-style sandbox. The sandbox runs a finite command without requiring a long-lived, health-checked canonical process.
The user-visible lifecycle should behave as follows:
Persistent, deployment-style sandboxes retain their existing canonical-main-process behavior, including health checks and the expectation that the main process remains running.
Acceptance Criteria
Alternatives Considered
Infer job semantics from any command supplied at creation
Treating every explicit command as a job would make finite commands convenient, but it would make long-running service commands ambiguous and could change the persistent canonical-main-process contract. An explicit lifecycle keeps intent clear.
Treat exit code 0 as success for every canonical main process
This would erase the distinction between a successfully completed job and a deployment-style sandbox whose required process unexpectedly stopped. The lifecycle mode should determine whether process exit represents completion or loss of service.
Create a retained sandbox, use
sandbox exec, then delete itThis is the current workaround. It requires multiple non-atomic operations, burdens callers with cleanup, and can leak retained resources when automation is interrupted.