Skip to content

feat(proxy): supervisor-proxied host-local endpoints — generalize the inference.local pattern for arbitrary host services #1633

Description

@maruiz93

Problem Statement

Sandbox agents increasingly need to reach services running on the host machine — build tools, repo provisioners, MCP tool servers, development utilities. Today there are two paths, and both have significant limitations.

Path 1: Bind to 0.0.0.0 + allowed_ips. The host service binds to all interfaces and the sandbox policy declares an allowed_ips CIDR matching the host IP that the container can reach. This works but has problems:

  • The service is exposed on all network interfaces — any process on any network can reach it, not just sandboxes. Token-based auth mitigates this but doesn't eliminate the attack surface.
  • The IP that containers use to reach the host varies by driver and networking mode. With rootless Podman + pasta, host.containers.internal resolves to a link-local address (169.254.1.2) that is handled by the pasta user-space proxy. With Docker, it's typically a bridge gateway (172.17.0.1). With Kubernetes, it requires explicit host_gateway_ip configuration.
  • On rootless Podman, the bridge gateway IP (e.g., 10.88.0.1) lives inside the container namespace and cannot be bound to from the host — bind() fails with EADDRNOTAVAIL. So binding to a specific internal interface is not an option; it's 0.0.0.0 or nothing.
  • allowed_ips must be templated with the correct IP at runtime, adding a fragile setup step that varies per platform. EDIT: allowed_ips is no longer required when endpoints are explicitly declared with host+port (PR fix(sandbox): trust exact declared private endpoints #1560). Validated on OpenShell v0.0.83 — see comment. The remaining motivation for this issue is the 0.0.0.0 exposure.
  • On macOS and WSL2, the bridge/gateway IP is often unreachable from the host entirely (bug: gateway crash-loops on macOS — binds to VM-internal podman bridge IP (10.89.0.1) on host #1358, bug(cluster): host.openshell.internal resolves to unreachable IP on Docker Desktop + WSL2 #811).

Path 2: iptables DNAT to loopback. Forward traffic from the container-reachable IP to 127.0.0.1 using NAT rules. This requires root, the route_localnet sysctl (system-wide security implications), and completely bypasses the supervisor proxy's L7 policy enforcement.

Meanwhile, OpenShell already solves a structurally identical problem for inference: inference.local lets sandbox agents reach host-side inference endpoints through the supervisor proxy, with TLS termination, credential handling, and L7 inspection. The service binds to 127.0.0.1 and the supervisor proxies the connection from inside the sandbox network namespace. This works identically across all drivers.

But inference.local is hardcoded for inference APIs — the intercepted hostname, the HTTP pattern matching, and the route bundle format are all inference-specific. There is no equivalent mechanism for a host service that exposes a build API, a repo provisioner, or an MCP tool server.

Concrete use cases

Host process model (GitHub Actions runners). We are running an experiment (fullsend-ai/experiments#28) where two host-side services — a container builder and a repo provisioner — provide capabilities to sandbox agents via REST APIs. The environment is Linux with rootless Podman + pasta networking on GitHub Actions workers.

Today the setup must:

  1. Bind both servers to 0.0.0.0 (can't bind to the bridge gateway IP 10.88.0.1 — it doesn't exist on the host in rootless mode)
  2. Template the host IP into allowed_ips in the sandbox policy EDIT: No longer needed — PR fix(sandbox): trust exact declared private endpoints #1560 removed this requirement for declared endpoints.
  3. Pass BUILDER_URL=http://host.openshell.internal:9090 to the sandbox
  4. Rely on bearer token auth as the only access control, since the services are exposed on all interfaces

If these servers could bind to 127.0.0.1 and the supervisor proxied the connection, the setup would be platform-independent, the servers would never be network-exposed, and L7 policy enforcement would apply.

Kubernetes sidecar model. The same API servers can be deployed as sidecar containers in the same pod as the OpenShell supervisor. Since Kubernetes mode always creates a nested network namespace (NetworkMode::Proxy is hardcoded in policy.rs:102-107), the agent cannot reach the sidecar at 127.0.0.1 directly — all traffic routes through the supervisor proxy via the veth pair (10.200.0.2 → 10.200.0.1:3128). The supervisor, running in the pod's root network namespace, can reach the sidecar at 127.0.0.1. The same host-local proxying mechanism applies: the proxy intercepts the request, evaluates L7 policy, and connects to 127.0.0.1:<port> in the pod netns where the sidecar is listening.

Both models use the same mechanism — 127.0.0.1 from the supervisor's perspective — which is what makes the proposal driver-agnostic.

Related work

Proposed Design

Generalize the inference.local supervisor-proxy mechanism so that policy-declared endpoints can be proxied through the supervisor to 127.0.0.1 on the host, instead of requiring the sandbox to reach the host service directly over the network.

Core mechanism

When the sandbox proxy receives a CONNECT request for a host-local endpoint:

  1. The proxy matches the destination against host-local endpoints declared in the active policy.
  2. Instead of resolving DNS and connecting to an external IP, the supervisor opens a TCP connection to 127.0.0.1:<port> on the host (the supervisor runs in the host/pod network namespace, not the sandbox netns).
  3. The proxy relays bytes between the sandbox client and the host-local upstream, applying L7 inspection and method/path rules from the policy.
  4. No new listeners are opened — this reuses the existing CONNECT proxy path.

This is structurally identical to how inference.local works today, minus the inference-specific pattern matching and route bundle format. The key difference is that the routing target (127.0.0.1:<port>) comes from the policy declaration rather than from a gateway-provided inference route bundle.

Policy surface — two options

The design question is how to declare host-local endpoints in the policy schema. Two options worth considering:

Option A: Reserved hostname. Introduce a new hostname (e.g., host.local or services.local) that the proxy intercepts, distinct from host.openshell.internal:

network_policies:
  builder:
    name: container-builder
    endpoints:
      - host: host.local
        port: 9090
        protocol: rest
        rules:
          - allow:
              method: POST
              path: /build
          - allow:
              method: GET
              path: /tools.json
  • Clean separation: host.openshell.internal = direct host access (existing, requires allowed_ips); host.local = supervisor-proxied host access (new, no allowed_ips needed).
  • Agents use http://host.local:9090/build — the hostname signals the routing path.
  • The proxy detects host.local like it detects inference.local — hardcoded interception, no DNS resolution needed.

Option B: Flag on existing endpoints. Add a host_local: true field to the endpoint schema:

network_policies:
  builder:
    name: container-builder
    endpoints:
      - host: host.openshell.internal
        port: 9090
        host_local: true
        protocol: rest
        rules:
          - allow:
              method: POST
              path: /build
  • Reuses host.openshell.internal — no new hostname.
  • The host_local flag changes proxy behavior: instead of connecting to the driver-injected IP, connect to 127.0.0.1.
  • Agent URLs stay http://host.openshell.internal:9090/build regardless of whether the endpoint is proxied or direct.

Both options have trade-offs. Option A makes the routing path visible in the URL; Option B keeps the URL stable but requires the proxy to check a flag before deciding how to connect.

What changes in the proxy

The generalizable components from the inference.local path:

  • CONNECT interception (proxy.rs:427): extend the hardcoded inference.local check to also match host-local endpoints from the active policy.
  • Connection dispatch: where inference.local proxies to the router, host-local endpoints connect directly to 127.0.0.1:<port>. No TLS termination, route bundle, or inference pattern matching needed — just TCP relay with L7 inspection.
  • SSRF enforcement: host-local endpoints bypass the normal SSRF tiers (they don't resolve DNS at all), similar to how inference.local bypasses OPA network policy at proxy.rs:374.
  • L7 rules: the existing method/path rule enforcement from the OPA engine applies unchanged.

What does NOT change

  • No new listeners (unlike PR fix(sandbox): add managed loopback proxy #1501).
  • No kernel-level changes (no iptables, no route_localnet).
  • No changes to host.openshell.internal DNS resolution or driver host-alias injection.
  • The existing allowed_ips path continues to work for users who prefer direct connectivity.

Driver compatibility

Because the supervisor always runs in the host/pod network namespace, 127.0.0.1 reaches co-located services from the supervisor's perspective. This works across drivers:

Driver Deployment model 127.0.0.1 reaches
Podman (rootless + pasta) Host process Host loopback — API servers run as host processes
Podman (rootful) Host process Host loopback
Docker (Linux) Host process Host loopback
Kubernetes Sidecar (same pod) Pod loopback — API servers run as sidecar containers sharing the pod netns
Docker Desktop (macOS/Windows) Host process VM loopback — host services need Docker Desktop's host-to-VM forwarding
VM (gvproxy) Host process VM loopback — could use 192.168.127.254 (gvproxy's host-loopback NAT) instead

Note: for Docker Desktop and VM drivers, "loopback" is the VM's loopback, not the physical host's. The design should consider whether the target address should be configurable (default 127.0.0.1, override via gateway config per driver).

Alternatives Considered

  1. Bind to 0.0.0.0 + allowed_ips (current approach). Works today but exposes services on all interfaces, varies by platform, and fails to bind to specific interfaces on rootless Podman. This is what we use now — it works, but the security and portability trade-offs motivate this proposal. EDIT: allowed_ips is no longer required (PR fix(sandbox): trust exact declared private endpoints #1560). The setup simplifies to 0.0.0.0 bind only. The security concern (all-interface exposure) remains and is the primary motivation for the host.local proposal.

  2. iptables DNAT to loopback. Requires root, route_localnet sysctl (system-wide — allows any network traffic to reach loopback on all interfaces), and bypasses the supervisor proxy entirely — no L7 policy enforcement. Not viable for production.

  3. socat/port-forwarding shim on the host. A user-space forwarder (e.g., socat TCP-LISTEN:9090,bind=$BRIDGE_GW,fork TCP:127.0.0.1:9090) per port. Extra process per service, still requires knowing the bridge IP, and provides no policy enforcement.

  4. Managed loopback proxy inside the sandbox netns (PR fix(sandbox): add managed loopback proxy #1501). Opens a new listener inside the sandbox network namespace that the supervisor dispatches. Closed by maintainer with the note: "would prefer not to open any additional listeners if we can avoid it" — but acknowledged as "likely enduring sandbox primitives." The current proposal avoids new listeners by reusing the existing CONNECT proxy path.

  5. Extend inference.local with more hardcoded hostnames. Add a second hardcoded hostname (e.g., tools.local) with a parallel interception path. Doesn't scale — each new use case would need another hardcoded hostname and dedicated routing logic. The policy-driven approach proposed here generalizes the mechanism.

  6. Separate Kubernetes pod with a Service. For Kubernetes deployments, the API server could run in its own pod with a cluster Service. This works with existing policy (host: api-server.ns.svc.cluster.local + allowed_ips) and doesn't need the host-local mechanism. However, it doesn't cover the host-process model (GitHub Actions, local dev) and adds deployment complexity compared to a sidecar.

Agent Investigation

Traced the inference.local implementation through the codebase to assess what's generalizable:

  • CONNECT interception (crates/openshell-sandbox/src/proxy.rs:427-451): hardcoded check for inference.local:443. The interception sends a 200 Connection Established response, then hands off to handle_inference_interception() for TLS termination and HTTP parsing. This is the insertion point for host-local endpoints.

  • L7 pattern matching (crates/openshell-sandbox/src/l7/inference.rs:62-83): detect_inference_pattern() matches HTTP method + path against a hardcoded list of OpenAI/Anthropic endpoints. The matching framework (method + path glob) is generic; only the pattern list is inference-specific. Host-local endpoints would not need this — L7 rules are already defined in the policy.

  • Route dispatching (crates/openshell-router/src/lib.rs:61-99): openshell_router::Router proxies matched requests to upstream endpoints via ResolvedRoute. This is the inference-specific routing layer — host-local endpoints would bypass it entirely and connect directly to 127.0.0.1:<port>.

  • SSRF enforcement tiers (proxy.rs:572-731): three tiers — trusted gateway (link-local only, for rootless Podman + pasta), allowed_ips (CIDR allowlist), and default reject. Loopback is always blocked in all tiers (is_always_blocked_ip() in crates/openshell-core/src/net.rs:45-62). Host-local endpoints would need a new tier or a bypass similar to how inference.local bypasses OPA at proxy.rs:374.

  • Policy schema (crates/openshell-policy/src/lib.rs:91-138): NetworkEndpointDef has generic fields — host, port, protocol (freeform string), rules (L7 method/path), allowed_ips. Adding host_local: bool (Option B) would be a one-field schema extension. Option A (reserved hostname) would require no schema changes — the proxy would match on hostname.

  • Kubernetes nested netns (crates/openshell-sandbox/src/policy.rs:102-107, sandbox/linux/netns.rs:61-186): Kubernetes always uses NetworkMode::Proxy, creating a nested network namespace with veth pair (10.200.0.1 ↔ 10.200.0.2). The supervisor runs in the pod's root netns and can reach sidecar containers at 127.0.0.1. The agent in the nested netns cannot — all traffic goes through the proxy. This confirms the sidecar model works with the same mechanism.

  • Validated experimentally: ran the fullsend host-side API server experiment on Linux with rootless Podman + pasta. Servers bound to 0.0.0.0:9090 and 0.0.0.0:9091, sandbox agent reached them via host.openshell.internal with allowed_ips policy. Both the builder (POST /build) and provisioner (POST /repo/provision) APIs completed successfully. The 0.0.0.0 bind is the workaround this proposal aims to eliminate. EDIT: Subsequently validated without allowed_ips — works with endpoint declaration alone (fullsend-ai/experiments#42).

Activity

  1. johntmyers commented on Jul 14, 2026

    @johntmyers
    Collaborator

    Currently, allowed_ips should not be required when the endpoint + port is explicitly declared. That was a change that was made in #1560 - have you validated this since opening this ticket?

    Based on that change, my understanding is that something like host.local that gets you to the actual host such that you can run arbitrary services on the host and restrict to listening on loopback is the primary ask. Is that right?

  2. maruiz93 commented on Jul 16, 2026

    @maruiz93
    Author

    Validated — allowed_ips is not required when the endpoint is explicitly declared with host+port. Tested on OpenShell v0.0.83, rootless Podman + pasta, Fedora 44:

    Endpoint Port Result
    builder (host.openshell.internal) 9090 200 OK
    provisioner (host.openshell.internal) 9091 200 OK
    example.com (not in policy) 80 Blocked

    We updated our experiment to drop allowed_ips and the HOST_IP templating entirely (fullsend-ai/experiments#42). The setup is significantly simpler now — no runtime IP resolution, no policy rendering step.

    And yes — the primary remaining ask is exactly that: a host.local mechanism so host services can bind to 127.0.0.1 instead of 0.0.0.0. The allowed_ips friction is solved, but the security surface isn't — services are still exposed on all interfaces. Being able to bind to loopback and have the supervisor proxy the connection (the same way inference.local works today) would close that gap.

  3. onoLA0808 commented on Aug 11, 2026

    @onoLA0808

    Adding another concrete use case for this feature: giving a sandboxed agent control of a real browser over the Chrome DevTools Protocol (CDP).

    Use case. Agent-driven browser automation (navigate, click, fill forms, screenshot) typically runs headful Chromium with --remote-debugging-port=9222 and connects the agent's tooling (Playwright/Puppeteer or a raw CDP client) to the DevTools endpoint. For a human-watchable setup, Chrome renders to Xvfb streamed over noVNC and the agent drives the same instance the operator watches.

    The catch: Chromium refuses to bind the DevTools port to anything but loopback — even with --remote-debugging-address=0.0.0.0 it binds 127.0.0.1:9222. So this is exactly the host-local-service problem described here: the service is only reachable on 127.0.0.1, which is always blocked from the sandbox.

    What I hit (matches this issue). OpenClaw under NemoClaw, Docker driver, WSL2. Agent netns 10.200.0.2, veth/supervisor 10.200.0.1, proxy 10.200.0.1:3128, Chrome on 127.0.0.1:9222 in the same container.

    Workarounds this issue already enumerates fail for the documented reasons:

    Node/socat forwarder 10.200.0.1:9222 → 127.0.0.1:9222 (Alternative #3). The agent's HTTP GET to /json/version succeeds through the proxy once a policy allows 10.200.0.1:9222 (ALLOWED ... [policy:browser_cdp engine:opa]). But the CDP WebSocket upgrade is denied:
    NET:OPEN [MED] DENIED -(0) -> 10.200.0.1:9222 [policy:- engine:opa]
    [reason:failed to resolve peer binary: No ESTABLISHED TCP connection found
    for port in /proc//net/tcp{,6}]

    The WebSocket is a separate raw TCP connection the proxy can't attribute to a binary at handshake time. Routed through the proxy via CONNECT it returns 403 instead — same root cause.

    Raw socket agent → 10.200.0.1:9222 → ECONNREFUSED (netns blocks it).
    The official OpenClaw browser plugin hits the equivalent wall — its gateway WebSocket drops (1006) with no policy log entry.

    Why this design would solve it. If a policy-declared host-local endpoint (Option A host.local:9222, or Option B host_local: true) let the supervisor proxy the CDP connection to 127.0.0.1:9222 including the WebSocket upgrade, with L7/policy applied, the agent could drive the browser through the sanctioned path with full openshell logs auditability — no 0.0.0.0 bind, no bridge-IP templating, no extra listeners.

    One design note: browser CDP is WebSocket-first — the /json/* HTTP endpoints just return a ws://.../devtools/... URL and all real traffic is the WS. So host-local proxying that only handles plain HTTP request/response wouldn't cover this; the relay needs to carry the WebSocket upgrade + bidirectional frames. That overlaps with the Slack Socket Mode WSS path noted as working in #1107, so the WS relay may be partly reusable.

    Happy to test a branch against this CDP setup — I have a reproducible environment (headful Chromium + Xvfb + noVNC + a CDP client) exercising exactly the loopback-WebSocket path. +1 on the host.local (Option A) direction.

  4. krishicks commented on Aug 14, 2026

    @krishicks
    Collaborator

    📋 triage-agent

    Removing state:triage-needed as workflow cleanup. Agent review of the issue's current state found substantive maintainer discussion, indicating that it is no longer awaiting initial triage. This label removal does not imply roadmap acceptance or implementation priority.

  5. LiamRandall commented on Aug 20, 2026

    @LiamRandall

    Adding a concrete integration use case + fresh data for this, from Cosmonic Desktop's OpenShell integration (a sandbox policy that makes a local WebAssembly runtime the execution outlet for a sandboxed coding agent; the agent talks to a host-local daemon via a stdio MCP server and verifies deployed apps on a loopback HTTP ingress).

    Setup: local single-user gateway (Homebrew service), Docker driver, macOS arm64, openshell 0.0.110. Sandbox policy declares network endpoints for the host dev daemon's ingress:

    endpoints:
      - { host: 127.0.0.1, port: 8200, enforcement: enforce }
      - { host: "*.localhost.cosmonic.sh", port: 8200, enforcement: enforce }  # wildcard DNS -> 127.0.0.1

    Observed on 0.0.110: the policy is accepted and effective, but any request to the allowlisted endpoints is hard-denied by the SSRF engine with no policy-level opt-in:

    $ curl http://<app>.localhost.cosmonic.sh:8200/   # inside the sandbox, /usr/bin/curl allowlisted
    {"detail":"GET <app>.localhost.cosmonic.sh:8200 blocked: internal address","error":"ssrf_denied"}
    

    (The literal 127.0.0.1 form never reaches the proxy at all; it's in the injected no_proxy — so container-loopback is the only interpretation available, and the DNS form is SSRF-blocked. Related: #2478, where the engine blocks even host.openshell.internal.)

    The SSRF guard is clearly right for shared/remote gateways. But for a local single-user gateway confining a local agent that must reach a host-local dev service, there is currently no sanctioned path at all; which is exactly the gap the supervisor-proxied inference.local-style host-local endpoint proposed here would close, without asking the host service to bind beyond loopback. Very interested in this landing; happy to validate against our policy/integration when there's something to test.

  6. github-actions commented on Sep 4, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

  7. maruiz93 commented on Sep 4, 2026

    @maruiz93
    Author

    I think this issues is still relevant

  8. github-actions commented on Sep 19, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

  9. maruiz93 commented on Sep 21, 2026

    @maruiz93
    Author

    remove stale

  10. maruiz93 commented on Sep 22, 2026

    @maruiz93
    Author

    @johntmyers may I ask more information about the not-planned thing? Are you open to contributions? Is it because you want to have a specific issue for addressing the specific case of host.local thing?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions