User Story
As an operator running OpenShell on Kubernetes with the Agent Sandbox v1beta1 API,
I want openshell sandbox stop followed by openshell sandbox start to return the sandbox to service,
so that I can pause and resume sandboxes without recreating them.
Problem Statement
After a stop→start cycle, a resumed sandbox never reaches Ready. Both the workload
pod and (for proxy-pod topology) the supervisor pod are Running 1/1, and the
Sandbox CR reports Ready=True (DependenciesReady), yet the gateway keeps the
sandbox at phase Starting indefinitely and sandbox start fails with a
300s readiness timeout.
The gateway derives phase from the Sandbox CR conditions. On the Agent Sandbox
v1beta1 API, the controller sets Suspended=True (PodTerminated) when a sandbox
is stopped and does not clear it when the sandbox is resumed. After start the
CR therefore carries both Ready=True and a stale Suspended=True. The gateway's
phase derivation treats a present Suspended=True condition as Stopped before
it ever consults the Ready condition, so a resumed-but-Ready sandbox is read as
not-running and the Starting lifecycle phase is never allowed to advance to
Ready.
This is independent of sandbox topology (it stems from the shared gateway phase
derivation) and is not present with the v1alpha1 API, which encodes desired state
as spec.replicas and does not emit a lingering Suspended condition.
Impact / Why This Matters
When this happens, openshell sandbox start blocks for 300s and then errors with
timed out after 300s waiting for sandbox to reach Ready, even though the
workload is actually running and healthy. The sandbox is stuck at Starting in
sandbox list/get and cannot be used through the gateway.
The only workaround is to delete and recreate the sandbox, which defeats the
purpose of stop/start (pausing to save resources while preserving the sandbox)
and discards any workspace/session state the operator intended to keep. Any
automation that relies on stop/start for cost control or scheduling is broken on
v1beta1 clusters.
Acceptance Criteria
Reproduction Steps
- Deploy an OpenShell gateway on a Kubernetes cluster using the Agent Sandbox
v1beta1 API (Agent Sandbox v0.5.0).
- Create a sandbox with a long-running entrypoint and wait for
Ready, e.g.
openshell sandbox create --name demo --driver-config-json '{"kubernetes":{"containers":{"agent":{"command":["sleep","infinity"]}}}}'
openshell sandbox stop demo and confirm the workload pod terminates.
openshell sandbox start demo.
- Observe: the workload pod returns to
Running 1/1, the CR shows
Ready=True (DependenciesReady) alongside a stale Suspended=True (PodTerminated),
but sandbox list shows phase Starting and the start command times out
after 300s.
Environment
- OpenShell: gateway built from
main (derive_phase unchanged through
fb6610df); observed with 0.0.112-dev gateway image.
- Runtime: Kubernetes with Agent Sandbox v1beta1 (v0.5.0); reproduced on
OpenShift 4.x with OVN-Kubernetes.
- Deployment: Helm chart, shared workspace mode. Observed with proxy-pod
topology but the phase derivation is topology-independent.
Logs
# CR after `sandbox start` — Ready but stale Suspended:
$ kubectl get sandbox default--demo -o jsonpath='{range .status.conditions[*]}{.type}={.status} ({.reason}){"\n"}{end}'
Ready=True (DependenciesReady)
Suspended=True (PodTerminated)
$ kubectl get sandbox default--demo -o jsonpath='spec.operatingMode={.spec.operatingMode}{"\n"}'
spec.operatingMode=Running
# CLI:
$ openshell sandbox start demo
Error: × timed out after 300s waiting for sandbox to reach Ready
# Code pointer (evidence, not a prescribed fix):
# crates/openshell-server/src/compute/mod.rs — derive_phase() returns
# SandboxPhase::Stopped on any Suspended=True condition before it checks Ready.
User Story
As an operator running OpenShell on Kubernetes with the Agent Sandbox v1beta1 API,
I want
openshell sandbox stopfollowed byopenshell sandbox startto return the sandbox to service,so that I can pause and resume sandboxes without recreating them.
Problem Statement
After a stop→start cycle, a resumed sandbox never reaches
Ready. Both the workloadpod and (for proxy-pod topology) the supervisor pod are
Running 1/1, and theSandbox CR reports
Ready=True (DependenciesReady), yet the gateway keeps thesandbox at phase
Startingindefinitely andsandbox startfails with a300s readiness timeout.
The gateway derives phase from the Sandbox CR conditions. On the Agent Sandbox
v1beta1 API, the controller sets
Suspended=True (PodTerminated)when a sandboxis stopped and does not clear it when the sandbox is resumed. After start the
CR therefore carries both
Ready=Trueand a staleSuspended=True. The gateway'sphase derivation treats a present
Suspended=Truecondition asStoppedbeforeit ever consults the
Readycondition, so a resumed-but-Ready sandbox is read asnot-running and the
Startinglifecycle phase is never allowed to advance toReady.This is independent of sandbox topology (it stems from the shared gateway phase
derivation) and is not present with the v1alpha1 API, which encodes desired state
as
spec.replicasand does not emit a lingeringSuspendedcondition.Impact / Why This Matters
When this happens,
openshell sandbox startblocks for 300s and then errors withtimed out after 300s waiting for sandbox to reach Ready, even though theworkload is actually running and healthy. The sandbox is stuck at
Startinginsandbox list/getand cannot be used through the gateway.The only workaround is to delete and recreate the sandbox, which defeats the
purpose of stop/start (pausing to save resources while preserving the sandbox)
and discards any workspace/session state the operator intended to keep. Any
automation that relies on stop/start for cost control or scheduling is broken on
v1beta1 clusters.
Acceptance Criteria
openshell sandbox stopthenopenshell sandbox startreturns the sandbox to phaseReady.openshell sandbox startcompletes without timing out when the workloadand supervisor pods are
Runningand the CR reportsReady=True.Suspended=Truecondition left on a resumed CR does not force thegateway phase to
Stopped/Startingwhen the sandbox is otherwise Readyand its desired operating mode is
Running.Reproduction Steps
v1beta1 API (Agent Sandbox v0.5.0).
Ready, e.g.openshell sandbox create --name demo --driver-config-json '{"kubernetes":{"containers":{"agent":{"command":["sleep","infinity"]}}}}'openshell sandbox stop demoand confirm the workload pod terminates.openshell sandbox start demo.Running 1/1, the CR showsReady=True (DependenciesReady)alongside a staleSuspended=True (PodTerminated),but
sandbox listshows phaseStartingand thestartcommand times outafter 300s.
Environment
main(derive_phaseunchanged throughfb6610df); observed with0.0.112-devgateway image.OpenShift 4.x with OVN-Kubernetes.
topology but the phase derivation is topology-independent.
Logs