Repository navigation
feat(proxy): supervisor-proxied host-local endpoints — generalize the inference.local pattern for arbitrary host services #1633
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on May 29, 2026 Currently,
allowed_ipsshould not be required when the endpoint + port is explicitly declared. That was a change that was made in #1560 - have you validated this since opening this ticket?Based on that change, my understanding is that something like
host.localthat gets you to the actual host such that you can run arbitrary services on the host and restrict to listening on loopback is the primary ask. Is that right?Reacted by Marta Añón RuizReacted by Marta Añón Ruiz- added 4 commits that reference this issue
on Jul 16, 2026 Validated —
allowed_ipsis not required when the endpoint is explicitly declared with host+port. Tested on OpenShell v0.0.83, rootless Podman + pasta, Fedora 44:Endpoint Port Result builder ( host.openshell.internal)9090 200 OKprovisioner ( host.openshell.internal)9091 200 OKexample.com(not in policy)80 Blocked We updated our experiment to drop
allowed_ipsand the HOST_IP templating entirely (fullsend-ai/experiments#42). The setup is significantly simpler now — no runtime IP resolution, no policy rendering step.And yes — the primary remaining ask is exactly that: a
host.localmechanism so host services can bind to127.0.0.1instead of0.0.0.0. Theallowed_ipsfriction is solved, but the security surface isn't — services are still exposed on all interfaces. Being able to bind to loopback and have the supervisor proxy the connection (the same wayinference.localworks today) would close that gap.Adding another concrete use case for this feature: giving a sandboxed agent control of a real browser over the Chrome DevTools Protocol (CDP).
Use case. Agent-driven browser automation (navigate, click, fill forms, screenshot) typically runs headful Chromium with --remote-debugging-port=9222 and connects the agent's tooling (Playwright/Puppeteer or a raw CDP client) to the DevTools endpoint. For a human-watchable setup, Chrome renders to Xvfb streamed over noVNC and the agent drives the same instance the operator watches.
The catch: Chromium refuses to bind the DevTools port to anything but loopback — even with --remote-debugging-address=0.0.0.0 it binds 127.0.0.1:9222. So this is exactly the host-local-service problem described here: the service is only reachable on 127.0.0.1, which is always blocked from the sandbox.
What I hit (matches this issue). OpenClaw under NemoClaw, Docker driver, WSL2. Agent netns 10.200.0.2, veth/supervisor 10.200.0.1, proxy 10.200.0.1:3128, Chrome on 127.0.0.1:9222 in the same container.
Workarounds this issue already enumerates fail for the documented reasons:
Node/socat forwarder 10.200.0.1:9222 → 127.0.0.1:9222 (Alternative #3). The agent's HTTP GET to /json/version succeeds through the proxy once a policy allows 10.200.0.1:9222 (ALLOWED ... [policy:browser_cdp engine:opa]). But the CDP WebSocket upgrade is denied:
NET:OPEN [MED] DENIED -(0) -> 10.200.0.1:9222 [policy:- engine:opa]
[reason:failed to resolve peer binary: No ESTABLISHED TCP connection found
for port in /proc//net/tcp{,6}]The WebSocket is a separate raw TCP connection the proxy can't attribute to a binary at handshake time. Routed through the proxy via CONNECT it returns 403 instead — same root cause.
Raw socket agent → 10.200.0.1:9222 → ECONNREFUSED (netns blocks it).
The official OpenClaw browser plugin hits the equivalent wall — its gateway WebSocket drops (1006) with no policy log entry.Why this design would solve it. If a policy-declared host-local endpoint (Option A host.local:9222, or Option B host_local: true) let the supervisor proxy the CDP connection to 127.0.0.1:9222 including the WebSocket upgrade, with L7/policy applied, the agent could drive the browser through the sanctioned path with full openshell logs auditability — no 0.0.0.0 bind, no bridge-IP templating, no extra listeners.
One design note: browser CDP is WebSocket-first — the /json/* HTTP endpoints just return a ws://.../devtools/... URL and all real traffic is the WS. So host-local proxying that only handles plain HTTP request/response wouldn't cover this; the relay needs to carry the WebSocket upgrade + bidirectional frames. That overlaps with the Slack Socket Mode WSS path noted as working in #1107, so the WS relay may be partly reusable.
Happy to test a branch against this CDP setup — I have a reproducible environment (headful Chromium + Xvfb + noVNC + a CDP client) exercising exactly the loopback-WebSocket path. +1 on the host.local (Option A) direction.
📋 triage-agent
Removing
state:triage-neededas workflow cleanup. Agent review of the issue's current state found substantive maintainer discussion, indicating that it is no longer awaiting initial triage. This label removal does not imply roadmap acceptance or implementation priority.- removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Aug 14, 2026 Adding a concrete integration use case + fresh data for this, from Cosmonic Desktop's OpenShell integration (a sandbox policy that makes a local WebAssembly runtime the execution outlet for a sandboxed coding agent; the agent talks to a host-local daemon via a stdio MCP server and verifies deployed apps on a loopback HTTP ingress).
Setup: local single-user gateway (Homebrew service), Docker driver, macOS arm64, openshell 0.0.110. Sandbox policy declares network endpoints for the host dev daemon's ingress:
endpoints: - { host: 127.0.0.1, port: 8200, enforcement: enforce } - { host: "*.localhost.cosmonic.sh", port: 8200, enforcement: enforce } # wildcard DNS -> 127.0.0.1
Observed on 0.0.110: the policy is accepted and effective, but any request to the allowlisted endpoints is hard-denied by the SSRF engine with no policy-level opt-in:
$ curl http://<app>.localhost.cosmonic.sh:8200/ # inside the sandbox, /usr/bin/curl allowlisted {"detail":"GET <app>.localhost.cosmonic.sh:8200 blocked: internal address","error":"ssrf_denied"}(The literal
127.0.0.1form never reaches the proxy at all; it's in the injectedno_proxy— so container-loopback is the only interpretation available, and the DNS form is SSRF-blocked. Related: #2478, where the engine blocks evenhost.openshell.internal.)The SSRF guard is clearly right for shared/remote gateways. But for a local single-user gateway confining a local agent that must reach a host-local dev service, there is currently no sanctioned path at all; which is exactly the gap the supervisor-proxied
inference.local-style host-local endpoint proposed here would close, without asking the host service to bind beyond loopback. Very interested in this landing; happy to validate against our policy/integration when there's something to test.This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Sep 4, 2026 I think this issues is still relevant
- removedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Sep 5, 2026 This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Sep 19, 2026 remove stale
- removedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Sep 21, 2026 @johntmyers may I ask more information about the not-planned thing? Are you open to contributions? Is it because you want to have a specific issue for addressing the specific case of
host.localthing?
Problem Statement
Sandbox agents increasingly need to reach services running on the host machine — build tools, repo provisioners, MCP tool servers, development utilities. Today there are two paths, and both have significant limitations.
Path 1: Bind to
0.0.0.0+allowed_ips. The host service binds to all interfaces and the sandbox policy declares anallowed_ipsCIDR matching the host IP that the container can reach. This works but has problems:host.containers.internalresolves to a link-local address (169.254.1.2) that is handled by the pasta user-space proxy. With Docker, it's typically a bridge gateway (172.17.0.1). With Kubernetes, it requires explicithost_gateway_ipconfiguration.10.88.0.1) lives inside the container namespace and cannot be bound to from the host —bind()fails withEADDRNOTAVAIL. So binding to a specific internal interface is not an option; it's0.0.0.0or nothing.EDIT:allowed_ipsmust be templated with the correct IP at runtime, adding a fragile setup step that varies per platform.allowed_ipsis no longer required when endpoints are explicitly declared with host+port (PR fix(sandbox): trust exact declared private endpoints #1560). Validated on OpenShell v0.0.83 — see comment. The remaining motivation for this issue is the0.0.0.0exposure.Path 2: iptables DNAT to loopback. Forward traffic from the container-reachable IP to
127.0.0.1using NAT rules. This requires root, theroute_localnetsysctl (system-wide security implications), and completely bypasses the supervisor proxy's L7 policy enforcement.Meanwhile, OpenShell already solves a structurally identical problem for inference:
inference.locallets sandbox agents reach host-side inference endpoints through the supervisor proxy, with TLS termination, credential handling, and L7 inspection. The service binds to127.0.0.1and the supervisor proxies the connection from inside the sandbox network namespace. This works identically across all drivers.But
inference.localis hardcoded for inference APIs — the intercepted hostname, the HTTP pattern matching, and the route bundle format are all inference-specific. There is no equivalent mechanism for a host service that exposes a build API, a repo provisioner, or an MCP tool server.Concrete use cases
Host process model (GitHub Actions runners). We are running an experiment (fullsend-ai/experiments#28) where two host-side services — a container builder and a repo provisioner — provide capabilities to sandbox agents via REST APIs. The environment is Linux with rootless Podman + pasta networking on GitHub Actions workers.
Today the setup must:
0.0.0.0(can't bind to the bridge gateway IP10.88.0.1— it doesn't exist on the host in rootless mode)Template the host IP intoEDIT: No longer needed — PR fix(sandbox): trust exact declared private endpoints #1560 removed this requirement for declared endpoints.allowed_ipsin the sandbox policyBUILDER_URL=http://host.openshell.internal:9090to the sandboxIf these servers could bind to
127.0.0.1and the supervisor proxied the connection, the setup would be platform-independent, the servers would never be network-exposed, and L7 policy enforcement would apply.Kubernetes sidecar model. The same API servers can be deployed as sidecar containers in the same pod as the OpenShell supervisor. Since Kubernetes mode always creates a nested network namespace (
NetworkMode::Proxyis hardcoded inpolicy.rs:102-107), the agent cannot reach the sidecar at127.0.0.1directly — all traffic routes through the supervisor proxy via the veth pair (10.200.0.2→10.200.0.1:3128). The supervisor, running in the pod's root network namespace, can reach the sidecar at127.0.0.1. The same host-local proxying mechanism applies: the proxy intercepts the request, evaluates L7 policy, and connects to127.0.0.1:<port>in the pod netns where the sidecar is listening.Both models use the same mechanism —
127.0.0.1from the supervisor's perspective — which is what makes the proposal driver-agnostic.Related work
allowed_ipsfor declared endpoints (would reduce friction on path 1, but doesn't solve the loopback or cross-platform problems)0.0.0.0→127.0.0.1→ configurable per driver)Proposed Design
Generalize the
inference.localsupervisor-proxy mechanism so that policy-declared endpoints can be proxied through the supervisor to127.0.0.1on the host, instead of requiring the sandbox to reach the host service directly over the network.Core mechanism
When the sandbox proxy receives a CONNECT request for a host-local endpoint:
127.0.0.1:<port>on the host (the supervisor runs in the host/pod network namespace, not the sandbox netns).This is structurally identical to how
inference.localworks today, minus the inference-specific pattern matching and route bundle format. The key difference is that the routing target (127.0.0.1:<port>) comes from the policy declaration rather than from a gateway-provided inference route bundle.Policy surface — two options
The design question is how to declare host-local endpoints in the policy schema. Two options worth considering:
Option A: Reserved hostname. Introduce a new hostname (e.g.,
host.localorservices.local) that the proxy intercepts, distinct fromhost.openshell.internal:host.openshell.internal= direct host access (existing, requiresallowed_ips);host.local= supervisor-proxied host access (new, noallowed_ipsneeded).http://host.local:9090/build— the hostname signals the routing path.host.locallike it detectsinference.local— hardcoded interception, no DNS resolution needed.Option B: Flag on existing endpoints. Add a
host_local: truefield to the endpoint schema:host.openshell.internal— no new hostname.host_localflag changes proxy behavior: instead of connecting to the driver-injected IP, connect to127.0.0.1.http://host.openshell.internal:9090/buildregardless of whether the endpoint is proxied or direct.Both options have trade-offs. Option A makes the routing path visible in the URL; Option B keeps the URL stable but requires the proxy to check a flag before deciding how to connect.
What changes in the proxy
The generalizable components from the
inference.localpath:proxy.rs:427): extend the hardcodedinference.localcheck to also match host-local endpoints from the active policy.inference.localproxies to the router, host-local endpoints connect directly to127.0.0.1:<port>. No TLS termination, route bundle, or inference pattern matching needed — just TCP relay with L7 inspection.inference.localbypasses OPA network policy atproxy.rs:374.method/pathrule enforcement from the OPA engine applies unchanged.What does NOT change
route_localnet).host.openshell.internalDNS resolution or driver host-alias injection.allowed_ipspath continues to work for users who prefer direct connectivity.Driver compatibility
Because the supervisor always runs in the host/pod network namespace,
127.0.0.1reaches co-located services from the supervisor's perspective. This works across drivers:127.0.0.1reaches192.168.127.254(gvproxy's host-loopback NAT) insteadNote: for Docker Desktop and VM drivers, "loopback" is the VM's loopback, not the physical host's. The design should consider whether the target address should be configurable (default
127.0.0.1, override via gateway config per driver).Alternatives Considered
Bind to
0.0.0.0+allowed_ips(current approach). Works today but exposes services on all interfaces, varies by platform, and fails to bind to specific interfaces on rootless Podman. This is what we use now — it works, but the security and portability trade-offs motivate this proposal. EDIT:allowed_ipsis no longer required (PR fix(sandbox): trust exact declared private endpoints #1560). The setup simplifies to0.0.0.0bind only. The security concern (all-interface exposure) remains and is the primary motivation for thehost.localproposal.iptables DNAT to loopback. Requires root,
route_localnetsysctl (system-wide — allows any network traffic to reach loopback on all interfaces), and bypasses the supervisor proxy entirely — no L7 policy enforcement. Not viable for production.socat/port-forwarding shim on the host. A user-space forwarder (e.g.,
socat TCP-LISTEN:9090,bind=$BRIDGE_GW,fork TCP:127.0.0.1:9090) per port. Extra process per service, still requires knowing the bridge IP, and provides no policy enforcement.Managed loopback proxy inside the sandbox netns (PR fix(sandbox): add managed loopback proxy #1501). Opens a new listener inside the sandbox network namespace that the supervisor dispatches. Closed by maintainer with the note: "would prefer not to open any additional listeners if we can avoid it" — but acknowledged as "likely enduring sandbox primitives." The current proposal avoids new listeners by reusing the existing CONNECT proxy path.
Extend
inference.localwith more hardcoded hostnames. Add a second hardcoded hostname (e.g.,tools.local) with a parallel interception path. Doesn't scale — each new use case would need another hardcoded hostname and dedicated routing logic. The policy-driven approach proposed here generalizes the mechanism.Separate Kubernetes pod with a Service. For Kubernetes deployments, the API server could run in its own pod with a cluster Service. This works with existing policy (
host: api-server.ns.svc.cluster.local+allowed_ips) and doesn't need the host-local mechanism. However, it doesn't cover the host-process model (GitHub Actions, local dev) and adds deployment complexity compared to a sidecar.Agent Investigation
Traced the
inference.localimplementation through the codebase to assess what's generalizable:CONNECT interception (
crates/openshell-sandbox/src/proxy.rs:427-451): hardcoded check forinference.local:443. The interception sends a200 Connection Establishedresponse, then hands off tohandle_inference_interception()for TLS termination and HTTP parsing. This is the insertion point for host-local endpoints.L7 pattern matching (
crates/openshell-sandbox/src/l7/inference.rs:62-83):detect_inference_pattern()matches HTTP method + path against a hardcoded list of OpenAI/Anthropic endpoints. The matching framework (method + path glob) is generic; only the pattern list is inference-specific. Host-local endpoints would not need this — L7 rules are already defined in the policy.Route dispatching (
crates/openshell-router/src/lib.rs:61-99):openshell_router::Routerproxies matched requests to upstream endpoints viaResolvedRoute. This is the inference-specific routing layer — host-local endpoints would bypass it entirely and connect directly to127.0.0.1:<port>.SSRF enforcement tiers (
proxy.rs:572-731): three tiers — trusted gateway (link-local only, for rootless Podman + pasta),allowed_ips(CIDR allowlist), and default reject. Loopback is always blocked in all tiers (is_always_blocked_ip()incrates/openshell-core/src/net.rs:45-62). Host-local endpoints would need a new tier or a bypass similar to howinference.localbypasses OPA atproxy.rs:374.Policy schema (
crates/openshell-policy/src/lib.rs:91-138):NetworkEndpointDefhas generic fields —host,port,protocol(freeform string),rules(L7 method/path),allowed_ips. Addinghost_local: bool(Option B) would be a one-field schema extension. Option A (reserved hostname) would require no schema changes — the proxy would match on hostname.Kubernetes nested netns (
crates/openshell-sandbox/src/policy.rs:102-107,sandbox/linux/netns.rs:61-186): Kubernetes always usesNetworkMode::Proxy, creating a nested network namespace with veth pair (10.200.0.1↔10.200.0.2). The supervisor runs in the pod's root netns and can reach sidecar containers at127.0.0.1. The agent in the nested netns cannot — all traffic goes through the proxy. This confirms the sidecar model works with the same mechanism.Validated experimentally: ran the fullsend host-side API server experiment on Linux with rootless Podman + pasta. Servers bound to
0.0.0.0:9090and0.0.0.0:9091, sandbox agent reached them viahost.openshell.internalwith. Both the builder (allowed_ipspolicyPOST /build) and provisioner (POST /repo/provision) APIs completed successfully. The0.0.0.0bind is the workaround this proposal aims to eliminate. EDIT: Subsequently validated withoutallowed_ips— works with endpoint declaration alone (fullsend-ai/experiments#42).