Repository navigation
feat: sandbox ingress — expose services from inside the network namespace #994
Description
Activity
Thanks for writing this up. I agree with the core problem statement: sandboxes need a first-class way to serve traffic, not only initiate outbound traffic. Web previews, webhook receivers, MCP servers, inter-agent APIs, and local service dependencies all need an ingress story.
My main piece of feedback is where the ingress boundary lives. I see ingress as a application/gateway-level OpenShell feature so the same model works across Kubernetes, Docker/Podman, local VM, and future compute drivers.
I have a mental model closer to something like Cloudflare Tunnels:
- The gateway is the client-facing/public ingress point.
- Sandboxes keep their supervisor-initiated connection to the gateway.
- The gateway authorizes incoming requests and multiplexes each accepted connection to the right sandbox over that supervisor session.
- The supervisor dials a validated target inside the sandbox, such as loopback TCP or a Unix socket.
- Compute drivers do not need to publish per-sandbox ports as the base data path.
We are actually actively working on this and should have some notes/prototypes we can share on it soon. I'm open to collaborating on the design and implementation here.
I'll detail how I've been thinking about it and we can compare notes and trade-offs.
Proposed direction
Use a sandbox-scoped service declaration model as the foundation:
openshell service create my-sandbox web --target-port 8080 --protocol http openshell service create my-sandbox mcp --target-port 3000 --protocol tcp openshell service list my-sandbox openshell service delete my-sandbox web
Those declarations should be OpenShell gateway resources, not Kubernetes Services. A declaration maps a stable service name to a validated target inside one sandbox:
(sandbox_id, service_name) -> target kind + target addressFor example:
(sbx_123, web) -> tcp_loopback:8080 (sbx_123, postgres) -> tcp_loopback:5432 (sbx_123, tools) -> unix_socket:/run/openshell/tools.sockThe gateway persists the declarations, tracks readiness based on the active supervisor session, and opens one relay stream per accepted client connection. This gives us service-specific authorization, auditing, lifecycle, and metrics without making the compute backend the routing authority. It also enabled much more dynamic routing without needing to patch Kubernetes specs.
Creating a service should also reconcile the sandbox policy automatically. Users should not have to remember to create a gateway service and then separately patch an ingress policy by hand. The service declaration is the user-facing intent; the gateway/control plane should materialize the corresponding policy on the sandbox.
Eg:
openshell service create my-sandbox web --target-port 8080 --protocol httpshould create or update a policy fragment equivalent to:
ingress: services: web: protocol: http target: kind: tcp_loopback port: 8080 via: gateway
The exact schema can change, but the ownership model is: service creation generates a managed policy entry, and service deletion removes that managed entry.
Gateway ingress and dynamic routing
We do want real ingress, not only local port forwarding. The dynamic external routing layer sits at the gateway.
For browser and HTTP-shaped services, use wildcard DNS and TLS/SNI/Host routing:
- A gateway operator configures a base domain, for example
open.sh. - Wildcard DNS points
*.open.shat the gateway or edge proxy. - The gateway allocates deterministic service hostnames from service declarations.
- Incoming TLS SNI and/or HTTP Host resolves to
(sandbox_id, service_name). - The gateway authenticates the request, opens a relay to that sandbox service, and proxies HTTP, WebSocket, SSE, or raw upgraded bytes.
The exact hostname shape is still a design decision, but the route could be derived from gateway-owned service metadata rather than Kubernetes object names.
Policy and security expectations
Ingress should be denied by default. A sandbox service should exist only when declared by an authorized actor or produced by an approved higher-level policy/profile flow.
Policy should remain the enforcement contract even though the gateway owns the route. Creating a service should automatically generate the policy needed for that service on the sandbox, and the supervisor should reject relay opens that are not present in the effective policy. The generated policy should allow the gateway relay for that specific service target; it should not imply general pod/network ingress to the sandbox. This gives us two checks:
- Gateway authorization decides whether the caller may access the named service.
- Sandbox policy decides whether the gateway may relay to the requested in-sandbox target.
- added a commit that references this issue
on May 1, 2026 Thanks for the writeup!
I want to flag a few use cases that I think are important and would like to discuss, because the way "gateway authorizes incoming requests" is currently described could rule them out.
Use cases
1. Per-sandbox authentication and authorization policies. Different sandboxes serve different audiences and need different auth schemes — sandbox A might require OIDC against tenant A's IdP, sandbox B might require mTLS with a tenant-specific CA. The gateway needs to support per-service (or at least per-sandbox) auth configuration: scheme, trust roots, IdP/issuer, claims requirements, secrets — all scoped so a misconfiguration or compromise in one sandbox's auth config can't affect another. A single gateway-wide auth config doesn't fit a multi-tenant deployment.
2. Calling a sandbox without gateway credentials. Several legitimate patterns require reaching a sandbox-exposed service without first authenticating to the gateway:
- Public APIs where the sandbox application itself decides who can call (e.g., an MCP server that gates by its own token, a preview server that's intentionally open).
- Custom or application-layer auth schemes the gateway doesn't natively understand.
For these, the gateway should support a per-service "transport-only" or "passthrough" mode: it validates that the route exists, applies transport-level protections (TLS, rate limit, optional IP allowlist), and forwards the request — including original headers — to the sandbox. The sandbox application performs caller authentication. This preserves the gateway's value as the routing/isolation layer without forcing it to be the auth layer for every service.
3. Exposing sandboxes directly via an external proxy (e.g., Envoy). In environments that already run an L7 proxy or service mesh, I'd want to use that proxy as the actual ingress point for sandboxes — Envoy terminates TLS (or does passthrough), runs JWT/ext_authz/OPA/WAF/rate-limiting, and forwards traffic to the sandbox without the OpenShell gateway being on the data path. The OpenShell gateway in this mode is a control plane: it owns service declarations, generates the corresponding sandbox ingress policy, and publishes route metadata that Envoy can consume (e.g., via an exported config, or a CRD the operator's existing toolchain already uses). The data path is Envoy → sandbox directly (via Service/Endpoint/EndpointSlice in K8s, or an equivalent target in other drivers).
This matters for several reasons:
- Scale. A single OpenShell gateway on the data path becomes a bottleneck at high sandbox counts and high per-service throughput. Letting an existing horizontally-scaled proxy fleet handle the bytes avoids reinventing data-plane scaling.
- Feature reuse. Envoy/Istio/Kong already have mature auth, observability, and traffic-management features. OpenShell needs to reimplement them, and in the best case scenario there is no need for two L7 proxies stacked back-to-back just to get the gateway's routing on top of Envoy's auth.
- Operational fit. Many deployments already standardize on a specific edge proxy. Sandbox ingress should plug into that, not require a parallel ingress path.
The gateway-tunnel model can stay the default for portability (Docker, VM, no-mesh K8s), but for K8s-with-mesh deployments there should be a "delegated data path" mode where the OpenShell gateway is purely a control plane and the proxy directly fronts sandboxes.
Let me know what you think!
- addedarea:gatewayGateway server and control-plane workGateway server and control-plane workarea:sandboxSandbox runtime and isolation workSandbox runtime and isolation workarea:cliCLI-related workCLI-related workarea:policyPolicy engine and policy lifecycle workPolicy engine and policy lifecycle work
on Jun 10, 2026
Problem Statement
The sandbox network namespace is egress-only. A process running inside the sandbox (an HTTP API, a webhook receiver, an MCP tool server, a development web server) is unreachable from outside the namespace. There is no mechanism — at the kernel, Kubernetes, or API level — for external traffic to reach a listening socket inside the sandbox.
This matters because agents increasingly need to serve, not just consume:
openshell forward, which only works for local access — other pods, load balancers, and CI runners cannot use the SSH tunnel.Current barriers
Five independent barriers prevent ingress into the sandbox:
net.ipv4.ip_forwardis never set on the host side of the veth pair. The kernel drops forwarded packets.10.200.0.2inside the namespace.containerPortentries. Standard Kubernetes service discovery does not work.Why SSH port forwarding is insufficient
openshell forwardtunnels traffic from the operator's local machine through the gateway's SSH connection. This solves local developer access but has fundamental limitations:Interaction with the split-pod proposal (#981)
Issue #981 proposes splitting the supervisor and agent into separate pods. In that architecture, the agent pod has its own CNI-assigned IP, so ingress becomes a standard Kubernetes Service pointing at the agent pod directly — no DNAT needed. However, the
InPod(single-pod) architecture remains the default and will coexist with the split-pod model. Ingress support for theInPodmode requires solving the DNAT/forwarding problem described here.The design should be structured so that the API surface (CLI flags, policy schema, proto fields) is shared across both modes, with only the enforcement mechanism differing (
InPoduses DNAT; split-pod uses native pod networking).Proposed Design
The design should address two questions: what ports to expose (the API surface) and how to forward traffic (the enforcement mechanism). This section focuses on problem constraints and design decisions rather than implementation details.
Where should ingress ports be declared?
There are three reasonable options:
Option A: In the sandbox spec (e.g.,
--publish 8080on create). The sandbox creator decides what to expose. Simple, Docker-like UX. But it means the sandbox creator can expose arbitrary ports without policy approval.Option B: In the policy (e.g.,
ingress.portsin the policy YAML). The policy author controls what can be exposed. Consistent with how egress is policy-controlled today. The sandbox creator cannot circumvent this.Option C: Both, via the provider profile layer model. Discussion #865 (tracked in #896) introduces a 3-layer policy composition model where provider profiles auto-inject endpoints into sandbox policy. Ingress ports could follow the same pattern:
openshell policy update --add-ingress 9090.This aligns ingress with the direction the project is already heading for egress policy. The provider profile already declares endpoints, binaries, and deny rules — adding an
ingress_portsfield is a natural extension.The decision between these options has implications for the proto schema (field on
SandboxTemplatevs.SandboxPolicy), the CLI surface, and whether ingress can be modified at runtime (policy supports hot reload; the sandbox spec does not).Kernel-level forwarding (InPod mode)
Regardless of where ports are declared, the
InPodmode needs iptables DNAT rules in the pod root namespace to forward traffic across the veth pair. The key constraints:/proc/sysis mounted read-only. The supervisor needs a fallback that temporarily remounts/proc/sysrw (requiresCAP_SYS_ADMIN, which the sandbox already has), writes the sysctl, and remounts back to ro. The remount is scoped to the pod's mount namespace, andnet.ipv4.ip_forwardis scoped to the pod's network namespace — neither affects the node or other pods.CAP_SYS_ADMINandCAP_NET_ADMINare already present.Kubernetes integration
sandbox-name.namespace.svc.cluster.local:port).Interaction with existing sandbox features
Alternatives Considered
1. HostPort mapping
Using Kubernetes
hostPortto expose the sandbox port on the node. Rejected: requires knowing the node IP (breaks service discovery), causes port conflicts across sandboxes on the same node, does not integrate with Kubernetes Services.2. Sidecar reverse proxy
Running envoy/socat as a sidecar that forwards into the sandbox namespace. Rejected: adds container overhead and image complexity, the sidecar runs outside sandbox policy enforcement, and the DNAT approach is simpler using existing veth infrastructure.
3. Modifying inner namespace routing for inbound
Adjusting the sandbox inner namespace to accept connections natively instead of DNAT in the root namespace. Rejected: the inner namespace intentionally has no default route back to the cluster (all egress goes through proxy), adding inbound routing would weaken the isolation model.
4. Defer entirely to the split-pod model (#981)
Wait for the split-pod architecture where ingress is trivial (agent pod has its own CNI IP). Rejected: the
InPodmode is the default and will remain so for Docker Desktop, local development, K3s clusters, and environments without gVisor. These deployments still need ingress support.Agent Investigation
Explored the codebase to assess feasibility:
crates/openshell-sandbox/src/sandbox/linux/netns.rs): Already has the veth pair, IP assignment, bypass detection iptables rules, andfind_iptables()helper. Ingress DNAT follows the same pattern asinstall_bypass_detection_rules().crates/openshell-driver-kubernetes/src/driver.rs):sandbox_template_to_k8s()already has env var injection andplatform_config_struct()for extracting typed config.proto/openshell.proto):SandboxTemplateusesgoogle.protobuf.Structforresourcesandvolume_claim_templates. Ingress could use the same pattern, or typed proto messages if the design goes with policy-level declaration.crates/openshell-server/src/compute/mod.rs):build_platform_config()already assembles a Struct from template fields.crates/openshell-sandbox/src/lib.rs): Port publication must happen after netns creation but before seccomp hardening. The insertion point is between the existing netns block and the seccomp prelude.crates/openshell-policy/): The policy language is entirely egress-oriented today. Adding aningresssection is structurally straightforward but would be the first inbound policy primitive.ingress_portsfield on provider profiles would let providers declare required inbound ports alongside their egress endpoints.Open questions for the maintainers
SandboxTemplate(creation-time, immutable), in the policy (hot-reloadable, policy-author-controlled), or auto-injected via provider profiles (Provider Enhancements -- Declarative Profiles, Auto-Injected Policy, Multi-Provider Inference #865)? This is the most consequential design decision.