Skip to content

bug(vm-driver): a sandbox that was never stopped ends in unrecoverable Error after a gateway restart #4209

Description

@ACodingfreak

User Story

I use OpenShell 0.1.2 with a local gateway and the VM compute driver (libkrun microVMs) on Ubuntu 24.04. I directly encountered this: after systemctl --user restart openshell-gateway, two of my three VM sandboxes ended in Error, and the only way out was to delete them.

This affects anyone who restarts or upgrades a local gateway that uses the VM driver: they can lose the sandbox and its /sandbox workspace.

Problem Statement

After the gateway restarts, a VM sandbox that has never been stopped and started since creation ends in Phase: Error, with this condition:

Ready: False (StartFailed) - Failed to start sandbox during gateway startup: VM sandbox is already running generation <64 hex characters>

The sandbox's VM is running again after the restart, but exec and SSH refuse the sandbox (sandbox '<name>' is not ready (phase: Error)). A sandbox that has been stopped and started at least once comes back Ready after the same restart.

Impact / Why This Matters

Any gateway restart does: systemctl --user restart openshell-gateway, a package upgrade, a gateway.toml change. A host reboot is expected to trigger it too.

  • There's no recovery path. openshell sandbox stop refuses (sandbox must be Ready to stop), and so does openshell sandbox start (sandbox must be Stopped, Completed, or a failed main-process Error to start). Later gateway restarts don't retry the sandbox. The only way out is sandbox delete, which discards /sandbox. Getting files back means copying the sandbox's overlay.ext4 and extracting from it by hand.
  • The workaround: stop and start every new VM sandbox once, right after creating it. That works (tested), but nothing in the CLI or docs tells users to do it.

Acceptance Criteria

  • A VM sandbox that was created and never stopped returns to Ready after a gateway restart, and openshell sandbox exec works in it.
  • A sandbox that a gateway restart has already put in Error with reason StartFailed can be recovered without deleting it and without losing /sandbox

Reproduction Steps

Prerequisites: the OpenShell 0.1.2 .deb, a local gateway with OPENSHELL_COMPUTE_DRIVER=vm running as the systemd user unit openshell-gateway, and KVM.

  1. Create two sandboxes and wait for both to be Ready:
    openshell sandbox create --name a --detach -- sleep infinity
    openshell sandbox create --name b --detach -- sleep infinity
    openshell sandbox list
  2. Stop and start only b:
    openshell sandbox stop b && openshell sandbox start b
  3. Restart the gateway:
    systemctl --user restart openshell-gateway
  4. After about 30 seconds, a is Error and b is Ready:
    openshell sandbox list
    openshell sandbox get a | grep StartFailed
  5. a can't be recovered:
    openshell sandbox stop a     # sandbox must be Ready to stop (current phase: Error)
    openshell sandbox start a    # sandbox must be Stopped, Completed, or a failed main-process Error to start

This also reproduces every time on a separate, throwaway gateway instance, for example a second VM-driver gateway on another bind_address with its own XDG_CONFIG_HOME and XDG_STATE_HOME, run under systemd-run --user and restarted with systemctl --user stop/start.

Suggested UX (if applicable)

No response

Environment

  • OpenShell: 0.1.2 (openshell --version: openshell 0.1.2; Debian package openshell 0.1.2-1).
  • OS: Ubuntu 24.04.4 LTS, kernel 6.8.0-139-generic, x86_64, KVM.
  • Deployment: local gateway installed from the .deb, running as the systemd user unit openshell-gateway (KillMode=control-group) on 127.0.0.1:17670. Compute driver vm (/usr/libexec/openshell/openshell-driver-vm, libkrun).

Logs

Gateway journal around the restart (`journalctl --user -u openshell-gateway`):


WARN openshell_server::compute: Failed to stop sandbox during gateway shutdown sandbox_id="fe76c66f-…" sandbox_name="vm1" error=code: 'The service is currently unavailable', …
WARN openshell_server::compute: Failed to start sandbox during gateway startup sandbox_id=fe76c66f-… sandbox_name=vm1 error=code: 'The system is not in a state required for the operation's execution', …
INFO openshell_server::compute: Started sandbox during gateway startup sandbox_id=… sandbox_name=dev phase=Provisioning recovered=false



$ openshell sandbox get vm1
  Phase: Error
  Conditions:
    - Ready: False (StartFailed) - Failed to start sandbox during gateway startup: VM sandbox is already running generation d925a7260e8c96f01c9c9f277c456f0a97316c83b567d479785042d4a03309c2
$ openshell sandbox stop vm1
Error:   × … message: "sandbox must be Ready to stop (current phase: Error)"
$ openshell sandbox start vm1
Error:   × … message: "sandbox must be Stopped, Completed, or a failed main-process Error to start (current phase: Error)"

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions