Repository navigation
feat: Support multiple sandboxes with different inference providers #776
Description
Activity
External operator data point for the design discussion: we ship a different multi-sandbox shape today that doesn't rely on a central inference router at all. Each sandbox has its own OpenClaw install with a distinct fallback chain (different primary model + fallback order per sandbox), and each sandbox's outbound network policy explicitly allows only the provider endpoints that sandbox actually uses. No
inference.localindirection — direct egress to upstream API endpoints, gated by per-sandbox policy.This works for us on
v0.0.25because the policy language lets us lock each sandbox down to its exact provider set (Slack + Firecrawl + Brave + the specific inference hosts the sandbox's chain needs). We swap model chains by editingopenclaw.jsonand restarting the gateway; sandbox-level isolation means one sandbox's provider choices can't leak into another's.Sharing so the "provider v2" design can account for both shapes:
- Centralized routing (the case described in the issue) — one
inference.local-style router per sandbox, multiple routers cluster-wide. - Direct per-sandbox egress — no central router, per-sandbox policies pointing straight at upstream endpoints.
Flag on the design: if provider v2 requires routing through a central component, it breaks case 2. If it's additive (operators can use a router when they need one, or stay on direct egress when they don't), both patterns keep working. Happy to share concrete per-sandbox policy snippets if useful.
- Centralized routing (the case described in the issue) — one
Hi @johntmyers - Any update on this issue.
Fresh real-world data point for this design discussion, plus a code-level map of what a multi-route design would need to touch (from tracing NVIDIA/NemoClaw#6315; verified on
mainabe42fb5and thev0.0.72tag).The need is still real, and the failure is now silent
NemoClaw#1248 (the issue this feature request was filed from) was closed after the hard-400 variant was fixed — but the underlying single-route contention is still there, and it now fails silently: NemoClaw#6315 (filed 2026-07-06, v0.0.74) shows that onboarding a second sandbox on a different provider/model silently re-points every existing sandbox's live inference within seconds, with no warning and no error. The failure mode moved from a loud 400 to a silent misroute, which is harder for users to notice and debug. NemoClaw currently works around the singleton by time-sharing the route — re-pointing it on every
connectwith a divergence warning — which cannot serve two always-on sandboxes concurrently.Current single-route enforcement points (what a multi-route design touches)
For whoever picks up the provider-v2 / multi-route design, these are the places where the one-user-route assumption is enforced today:
- Route-name whitelist —
crates/openshell-server/src/inference.rs::effective_route_nameaccepts exactlyinference.localandsandbox-system; everything else is rejected withunknown route_name. This is the hard gate:SetClusterInference/GetClusterInferencecan already carry arbitraryroute_namestrings in the proto (proto/inference.proto), so the wire format doesn't need to change — the whitelist and the store keying do. - Bundle resolution —
resolve_inference_bundle(same file) returns the two fixed routes plus a content-hash revision. A multi-route design needs either per-sandbox filtering here (the handler already authenticates a sandbox principal, so the sandbox identity is available at resolution time) or a route-selection field in the bundle. - Sandbox-side refresh —
crates/openshell-supervisor-network/src/inference_routes.rsfetches the bundle at startup and refreshes every 5 s in cluster mode (DEFAULT_ROUTE_REFRESH_INTERVAL_SECS). This is why a route flip propagates to running sandboxes within seconds. No changes needed here if the bundle arrives pre-filtered per sandbox. - Router pinning —
crates/openshell-router/src/backend.rsdeliberately rewrites the request's JSONmodelfield to the route's configured model and injects the route's credentials/base URL ("so a sandbox cannot pick a different upstream model than whatinference setconfigured"). This operator-pinning property is worth preserving in a multi-route design — the route stays operator-configured, it just becomes selected per sandbox principal instead of shared.
Given that the sandbox principal is already authenticated at
GetInferenceBundletime, the smallest-surface design seems to be: key user routes by sandbox (or by an operator-assigned route name mapped to sandboxes), filter inresolve_inference_bundle, and keepinference.localas the in-sandbox addressing name — no sandbox-side or policy changes required. Happy to turn this into a fuller proposal or a spike if maintainers want to revive this.- Route-name whitelist —
This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 30, 2026 📋 triage-agent
Closing the managed multi-route design as superseded by #3195. The user outcome now uses a different architecture: each sandbox can attach its own provider profile and call that provider native endpoint directly.
inference.local, cluster-wide inference routes, and numbered local route hostnames were intentionally removed rather than extended.
Problem Statement
NemoClaw has a requirement to support multiple sandboxes with different inference provided (for example: OpenAI and Antropic))
See NemoClaw issue: NVIDIA/NemoClaw#1248
NemoClaw policy requires to direct the inferences via the inference router (inference.local). That means, provider endpoints should not be allowed to be accessed directly via relaxing the networking policies.
Proposed Design
OpenShell team is looking into the "provider v2" system, as discussed over slack.
OpenShell inference route is a cluster-wide singleton. OpenShell should support multiple inference routes (i.e. inference1.local, inference2.local and so on).
Alternatives Considered
NemoClaw policy requires to direct the inferences via the inference router (inference.local). That means, provider endpoints should not be allowed to be accessed directly via relaxing the networking policies.
So, this is not a feasible alternative.
Agent Investigation
No response
Checklist