Repository navigation
[Bug] L7 egress proxy denies all CONNECT requests on Docker Desktop + WSL2 (amd64) #681
Description
Activity
For anyone interested in the patch while I work through the vouch system, you can see the pr in my fork.
- added a commit that references this issue
on Mar 30, 2026 I am curious why this hasn't been encountered by other WSL users. Any idea why?
my best guess is a combination of factors.
this issue only appears when network policies include a
binariessection. that setting causes the proxy to resolve the peer connection to a specific binary by looking at/proc/<pid>/net/tcp. on wsl2 with docker desktop, that code path doesn't work as expected because connections redirected via iptables/dnat don't show up in the per-pid table (they only appear in the global/proc/net/tcp).without per-binary enforcement the identity resolution step isn't needed, so the problem doesn't surface. additionally, the failure is fairly quiet — the opa deny is logged at
info!()level while the sandbox defaults towarn, so it only shows a generic 403 with no clear diagnostic information.this might explain why it hasn't come up more often.
Ok thanks that is helpful. I will take a look at this more thoroughly tomorrow. We don't have WSL setup for CI yet but I am ok getting this in at our current stage of development if you are able to verify it working.
@johntmyers – following up on your comment about verifying the fix.
i ran some reproduction tests on my wsl2 setup (windows 11, docker desktop 29.3.1, wsl2 kernel 6.6.87.2, openshell v0.0.16). i enabled
rust_log=openshell_sandbox=debugand ran everything inside the sandbox network namespace.what i found:
- with a properly configured policy that includes
nvidia_web(allowingcurltowww.nvidia.com:443), identity resolution works on the stock binary. the proxy successfully resolves the binary and opa allows the request. - the 403s i was originally seeing were not caused by a failure in
parse_proc_net_tcp(). they were caused by the active gateway policy not including the expected rules (the baked-in/etc/openshell/policy.yamlwas being overridden by the gateway's own policy set).
once i added the missing policy via
openshell policy set, everything worked as expected with the stock binary.observability note:
the sandbox defaults to
warnlog level, but the connect allow/deny decision is only logged atinfo. this made the denials completely silent at the default level, which led me to assume a deeper problem in identity resolution. switching to debug level immediately showed that resolution was succeeding and the deny was purely policy-related. it might be worth considering logging connect denials atwarnso they're visible by default.regarding pr #684:
i also tested the patched binary (with the
/proc/net/tcpglobal fallback). on my current setup, both/proc/1/net/tcpand/proc/net/tcpcontain the same entries, so the fallback path never triggered.i was unable to reproduce the per-pid vs global
/proc/net/tcpdiscrepancy i originally reported. it's possible a recent docker desktop or wsl2 kernel update changed the behavior, or that my original 403s were always policy-related and i misattributed them.recommendation:
i believe that the fallback added in pr #684 is defensive and harmless — it tries the per-pid path first and falls back to global if needed. however, since i can't demonstrate it fixing a real issue on my setup anymore, i'll leave it to you whether it's worth keeping as a resilience measure or if you'd prefer to close the pr.
happy to provide any additional logs or run further tests if helpful.
- with a properly configured policy that includes
Thank you for the update. I've opened #704 to address the observability issue and will create a PR. At this point I am going to close out #684. My concern with falling back to the global table is that we could accidentally hit a port collision because this table contains connections from all network namespaces and all processes on the host. So it very well could return an inode matching the port but to a process outside of the sandbox. One mitigation is to match on both local port and remote port which can reduce collisions but it's not impossible. So I'd argue while the fallback makes sense it's not worth the security risk without much deeper exploration which is hard to justify now.
Reacted by David Peden@davidpeden3 which logs were you looking at? The actual logs emitted from OpenShell or container logs from Kube? From what I can tell the denial logs should be available via the sandbox itself with
openshell sandbox logs ...or via the TUI which you can launch withopenshell term? If I'm understanding this right we'd need to update the logging so the same logs come out via container logs. But I do want to make sure you're aware of the logs we emit directly from the sandbox to the gateway.Hi @davidpeden3 — I hit the same class of bug in a different scenario: supervisor helpers live in pid 1's network namespace, not the workload's, so the policy proxy could not find their CONNECT peers in
/proc/<entrypoint_pid>/net/tcp.I opened #956 with a narrower fallback than #684: it checks
/proc/1/net/tcp, not the init-namespace global/proc/net/tcp.That should avoid the port-collision concern, and I believe it should also cover this WSL2 case. Pid 1 inside the sandbox pod is in the pod's network namespace, so iptables-REDIRECT'd connections should show up in
/proc/1/net/tcptoo.I do not have a WSL2 setup to confirm. If you still have the repro, mind trying #956? If it works, it should close this issue with a narrower trust boundary than the previous approach.
@Kh4L — tested #956 on my WSL2 setup (Windows 11, Docker Desktop 29.3.1, WSL2 kernel 6.6.87.2). the fix works.
setup:
PID 1 (supervisor) and the workload (PID 71,
sleep infinity) are in different network namespaces:/proc/1/ns/net -> net:[4026532836] /proc/71/ns/net -> net:[4026532921]/proc/71/net/tcp(workload) is empty./proc/1/net/tcp(supervisor) has the ESTABLISHED connections.A/B results:
binary CONNECT integrate.api.nvidia.com:443 result stock v0.0.23 HTTP/1.1 403 Forbiddenpeer resolution fails — workload netns table is empty PR #956 HTTP/1.1 200 Connection Establishedfallback to /proc/1/net/tcpresolves the peertested with a policy that scopes
binaries: [{ path: /usr/bin/curl }]on the nvidia endpoint. curl was run as the sandbox user from the workload network namespace viansenter -t 71 --net -S 1000 -G 1000.this also retroactively explains the original #681 report — the per-pid vs global discrepancy i saw back in march was real, i just couldn't reproduce it later because my subsequent tests were running from the supervisor namespace without realizing it.
@davidpeden3 awesome!
@johntmyers can I get a review on #956 ?
Agent Diagnostic
crates/openshell-sandbox/src/)proxy.rsto trace the CONNECT handling flow: request parsing (line 301) →evaluate_opa_tcp()(line 340) → OPA deny → 403 response (line 414)procfs.rsto understand binary identity resolution:resolve_tcp_peer_identity()→parse_proc_net_tcp()→find_pid_by_socket_inode()parse_proc_net_tcp()only reads/proc/<pid>/net/tcp{,6}(line 178)/proc/1/net/tcpcontents between macOS Docker (arm64) and WSL2 Docker (amd64) — sandbox user connections visible on macOS, absent on WSL2/proc/net/tcp(global view) on WSL2 — sandbox user connections ARE visible there/proc/<pid>/net/tcp, only in global/proc/net/tcp/proc/net/tcp{,6}as fallback inparse_proc_net_tcp(). Built patched binary on amd64 and verified the fix.Description
The sandbox egress proxy at
10.200.0.1:3128returns HTTP 403 Forbidden for allHTTP CONNECTrequests when running on Docker Desktop with WSL2 (Windows 11, amd64). The same container image, policy, and binary work correctly on Docker Desktop with Apple Hypervisor (macOS, arm64).parse_proc_net_tcp()inprocfs.rsonly reads/proc/<entrypoint_pid>/net/tcp{,6}to resolve the socket inode for the peer connection. On WSL2/Docker Desktop, connections from the sandbox user namespace that are redirected via iptables REDIRECT/DNAT to the proxy are not visible in the per-PID table — they only appear in the global/proc/net/tcp.This causes
resolve_tcp_peer_identity()to fail, the proxy cannot identify the calling binary, no network policy matches, and OPA denies the request.Forward proxy requests (
GET http://...) are also affected since they use the sameevaluate_opa_tcp()code path.Reproduction Steps
socatandcurlto connect to a non-loopback IP:curl -v -p -x http://10.200.0.1:3128 http://<gateway-ip>:18789/HTTP/1.1 403 Forbidden(CONNECT tunnel failed)HTTP/1.1 200 Connection EstablishedDiagnostic from inside the pod on WSL2:
Environment
ghcr.io/nvidia/openshell/cluster:0.0.16)Logs
Agent-First Checklist
debug-openshell-cluster,debug-inference,openshell-cli)