Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions architecture/compute-runtimes.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,11 @@ whose compute resource is still provisioning without exposing contradictory publ
readiness signals. When the driver reports runtime readiness, its ready condition
is published without waiting for a supervisor session.

The supervisor keeps retrying session establishment while the gateway is unavailable.
Its control readiness socket remains absent until the gateway accepts a session and
is removed if that session disconnects. A transient gateway delay during startup
therefore leaves the sandbox provisioning without terminating the supervisor.

**Session precedence over lagging driver snapshots:** A supervisor session can only be
established by a running workload. When `set_supervisor_session_state` promotes the
store record to `Ready` on session connect, a driver watch event may still arrive
Expand Down
25 changes: 5 additions & 20 deletions crates/openshell-supervisor-process/src/delegated.rs
Original file line number Diff line number Diff line change
Expand Up @@ -156,7 +156,7 @@ pub async fn start_boundary_access(

let (session_task, session_readiness) = match (openshell_endpoint, sandbox_id) {
(Some(endpoint), Some(id)) => {
let (task, mut accepted) = crate::supervisor_session::spawn_with_readiness(
let (task, accepted) = crate::supervisor_session::spawn_with_readiness(
endpoint.to_string(),
id.to_string(),
ssh_socket_path,
Expand All @@ -168,25 +168,10 @@ pub async fn start_boundary_access(
session_id_updates: supervisor_session_updates,
},
);
let accepted_result =
tokio::time::timeout(Duration::from_secs(10), accepted.wait_for(|ready| *ready))
.await
.map(|result| result.map(|_| ()));
match accepted_result {
Ok(Ok(())) => (Some(task), Some(accepted)),
Ok(Err(_)) => {
task.abort();
return Err(miette::miette!(
"supervisor session ended before gateway acceptance"
));
}
Err(_) => {
task.abort();
return Err(miette::miette!(
"gateway did not accept supervisor session within 10 seconds"
));
}
}
// Session establishment retries through gateway restarts. The
// readiness socket remains absent until the gateway accepts the
// session, so a transient delay cannot kill the supervisor.
(Some(task), Some(accepted))
}
_ => (None, None),
};
Expand Down
36 changes: 24 additions & 12 deletions crates/openshell-supervisor/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -100,21 +100,24 @@ impl ControlReadiness {
path: std::path::PathBuf,
mut session_readiness: Option<tokio::sync::watch::Receiver<bool>>,
) -> Result<Self> {
if session_readiness
prepare_control_readiness_path(&path)?;
let listener = if session_readiness
.as_ref()
.is_some_and(|readiness| !*readiness.borrow())
{
return Err(miette::miette!(
"supervisor session is not ready when starting health listener"
));
}
prepare_control_readiness_path(&path)?;
let listener = tokio::net::UnixListener::bind(&path)
.into_diagnostic()
.wrap_err_with(|| format!("bind supervisor readiness socket on {}", path.display()))?;
None
} else {
Some(
tokio::net::UnixListener::bind(&path)
.into_diagnostic()
.wrap_err_with(|| {
format!("bind supervisor readiness socket on {}", path.display())
})?,
)
};
let task_path = path.clone();
let task = tokio::spawn(async move {
let mut listener = Some(listener);
let mut listener = listener;
loop {
let session_unready = session_readiness
.as_ref()
Expand Down Expand Up @@ -4754,10 +4757,19 @@ mod tests {
async fn control_readiness_tracks_supervisor_session() {
let root = tempfile::tempdir().unwrap();
let path = root.path().join("health.sock");
let (session_tx, session_rx) = tokio::sync::watch::channel(true);
let (session_tx, session_rx) = tokio::sync::watch::channel(false);
let _readiness = ControlReadiness::start(path.clone(), Some(session_rx))
.expect("start readiness listener");
check_control_readiness(&path).expect("accepted session is ready");
assert!(check_control_readiness(&path).is_err());

session_tx.send_replace(true);
timeout(Duration::from_secs(1), async {
while check_control_readiness(&path).is_err() {
tokio::task::yield_now().await;
}
})
.await
.expect("first accepted session creates readiness socket");

session_tx.send_replace(false);
timeout(Duration::from_secs(1), async {
Expand Down
2 changes: 2 additions & 0 deletions docs/how-it-works/sandboxes/runtimes.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,8 @@ The driver looks up sandbox images in local Docker or Podman before pulling from

VM sandboxes have no network interface. All traffic flows through the OpenShell supervisor on the host. For networks that require a corporate proxy, the VM driver accepts the same `https_proxy` and `proxy_*` options as Podman. To reach a proxy on the gateway host, use `http://host.openshell.internal:<port>`.

If the gateway is briefly unavailable while a VM starts, its supervisor retries the session connection. The sandbox stays in provisioning until the gateway accepts the session.

## Kubernetes Driver

The Kubernetes driver runs sandboxes as pods in a sandbox namespace. It requires the [Agent Sandbox](https://github.com/kubernetes-sigs/agent-sandbox) controller. Install the gateway with the Helm chart; see [Kubernetes setup](/kubernetes/setup).
Expand Down
Loading