Skip to content

feat: Support multiple sandboxes with different inference providers #776

Description

@paritoshd-nv

Problem Statement

NemoClaw has a requirement to support multiple sandboxes with different inference provided (for example: OpenAI and Antropic))

See NemoClaw issue: NVIDIA/NemoClaw#1248

NemoClaw policy requires to direct the inferences via the inference router (inference.local). That means, provider endpoints should not be allowed to be accessed directly via relaxing the networking policies.

Proposed Design

OpenShell team is looking into the "provider v2" system, as discussed over slack.

OpenShell inference route is a cluster-wide singleton. OpenShell should support multiple inference routes (i.e. inference1.local, inference2.local and so on).

Alternatives Considered

NemoClaw policy requires to direct the inferences via the inference router (inference.local). That means, provider endpoints should not be allowed to be accessed directly via relaxing the networking policies.
So, this is not a feasible alternative.

Agent Investigation

No response

Checklist

  • I've reviewed existing issues and the architecture docs
  • This is a design proposal, not a "please build this" request

Activity

  1. mjamiv commented on Apr 13, 2026

    @mjamiv
    Contributor

    External operator data point for the design discussion: we ship a different multi-sandbox shape today that doesn't rely on a central inference router at all. Each sandbox has its own OpenClaw install with a distinct fallback chain (different primary model + fallback order per sandbox), and each sandbox's outbound network policy explicitly allows only the provider endpoints that sandbox actually uses. No inference.local indirection — direct egress to upstream API endpoints, gated by per-sandbox policy.

    This works for us on v0.0.25 because the policy language lets us lock each sandbox down to its exact provider set (Slack + Firecrawl + Brave + the specific inference hosts the sandbox's chain needs). We swap model chains by editing openclaw.json and restarting the gateway; sandbox-level isolation means one sandbox's provider choices can't leak into another's.

    Sharing so the "provider v2" design can account for both shapes:

    1. Centralized routing (the case described in the issue) — one inference.local-style router per sandbox, multiple routers cluster-wide.
    2. Direct per-sandbox egress — no central router, per-sandbox policies pointing straight at upstream endpoints.

    Flag on the design: if provider v2 requires routing through a central component, it breaks case 2. If it's additive (operators can use a router when they need one, or stay on direct egress when they don't), both patterns keep working. Happy to share concrete per-sandbox policy snippets if useful.

  2. paritoshd-nv commented on Apr 23, 2026

    @paritoshd-nv
    Author

    Hi @johntmyers - Any update on this issue.

  3. Hokonoken commented on Jul 7, 2026

    @Hokonoken

    Fresh real-world data point for this design discussion, plus a code-level map of what a multi-route design would need to touch (from tracing NVIDIA/NemoClaw#6315; verified on main abe42fb5 and the v0.0.72 tag).

    The need is still real, and the failure is now silent

    NemoClaw#1248 (the issue this feature request was filed from) was closed after the hard-400 variant was fixed — but the underlying single-route contention is still there, and it now fails silently: NemoClaw#6315 (filed 2026-07-06, v0.0.74) shows that onboarding a second sandbox on a different provider/model silently re-points every existing sandbox's live inference within seconds, with no warning and no error. The failure mode moved from a loud 400 to a silent misroute, which is harder for users to notice and debug. NemoClaw currently works around the singleton by time-sharing the route — re-pointing it on every connect with a divergence warning — which cannot serve two always-on sandboxes concurrently.

    Current single-route enforcement points (what a multi-route design touches)

    For whoever picks up the provider-v2 / multi-route design, these are the places where the one-user-route assumption is enforced today:

    1. Route-name whitelist — crates/openshell-server/src/inference.rs::effective_route_name accepts exactly inference.local and sandbox-system; everything else is rejected with unknown route_name. This is the hard gate: SetClusterInference / GetClusterInference can already carry arbitrary route_name strings in the proto (proto/inference.proto), so the wire format doesn't need to change — the whitelist and the store keying do.
    2. Bundle resolution — resolve_inference_bundle (same file) returns the two fixed routes plus a content-hash revision. A multi-route design needs either per-sandbox filtering here (the handler already authenticates a sandbox principal, so the sandbox identity is available at resolution time) or a route-selection field in the bundle.
    3. Sandbox-side refresh — crates/openshell-supervisor-network/src/inference_routes.rs fetches the bundle at startup and refreshes every 5 s in cluster mode (DEFAULT_ROUTE_REFRESH_INTERVAL_SECS). This is why a route flip propagates to running sandboxes within seconds. No changes needed here if the bundle arrives pre-filtered per sandbox.
    4. Router pinning — crates/openshell-router/src/backend.rs deliberately rewrites the request's JSON model field to the route's configured model and injects the route's credentials/base URL ("so a sandbox cannot pick a different upstream model than what inference set configured"). This operator-pinning property is worth preserving in a multi-route design — the route stays operator-configured, it just becomes selected per sandbox principal instead of shared.

    Given that the sandbox principal is already authenticated at GetInferenceBundle time, the smallest-surface design seems to be: key user routes by sandbox (or by an operator-assigned route name mapped to sandboxes), filter in resolve_inference_bundle, and keep inference.local as the in-sandbox addressing name — no sandbox-side or policy changes required. Happy to turn this into a fuller proposal or a spike if maintainers want to revive this.

  4. github-actions commented on Aug 30, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

  5. johntmyers commented on Sep 17, 2026

    @johntmyers
    Collaborator

    📋 triage-agent

    Closing the managed multi-route design as superseded by #3195. The user outcome now uses a different architecture: each sandbox can attach its own provider profile and call that provider native endpoint directly. inference.local, cluster-wide inference routes, and numbered local route hostnames were intentionally removed rather than extended.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:staleInactive item at risk of automatic closure.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions