Repository navigation
bug: GPU passthrough fails on WSL2 — NVML init fails without CDI mode and libdxcore.so #404
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Mar 17, 2026 This issue appears to have been opened without an agent investigation.
OpenShell is an agent-first project - please point your coding agent at the repo and have it diagnose this before we triage. Your agent can load skills like
debug-openshell-cluster,debug-inference,openshell-cli, andgenerate-sandbox-policy.See CONTRIBUTING.md for the full workflow.
I'm the human filing these issues, feel free to direct questions and/or rage my way.
Relates to setting up Nvidia NemoClaw issue NVIDIA/NemoClaw#208- added 2 commits that reference this issue
on Mar 17, 2026 - addedtopic:compatibilityCompatibility-related workCompatibility-related work
on Mar 17, 2026 Fix in #411
- added a commit that references this issue
on Mar 17, 2026 Confirmed the full CDI pipeline fix working on RTX 5090 Laptop (24GB), WSL2 Ubuntu 24.04, Docker Desktop, OpenShell 0.0.7. nvidia-smi fully functional inside a GPU-enabled sandbox.
A few additional findings beyond the original writeup:
- CDI spec needs a UUID device entry, not just name: all. The device plugin allocates GPUs by UUID (e.g. nvidia.com/gpu=GPU-a0edb6f7-...), but nvidia-ctk cdi generate in WSL mode only creates an all entry. Without the UUID entry, containerd returns unresolvable CDI devices.
- containerd needs enable_cdi = true in the k3s config. Switching the nvidia runtime to CDI mode isn't enough — containerd itself must have CDI enabled via /var/lib/rancher/k3s/agent/etc/containerd/config.toml.tmpl.
- containerd must be restarted after CDI/runtime changes. The k3s-embedded containerd won't pick up the new spec or runtime mode until its process is killed and respawned.
- nvidia pods must be force-deleted after the restart so they reschedule with the new CDI config. Existing CrashLoopBackOff pods keep failing against the old config.
- sed append breaks CDI YAML structure — used awk with a temp file to patch the spec safely.
Automated script: https://github.com/thenewguardai/tng-nemoclaw-quickstart/blob/main/scripts/wsl2-gpu-deploy.sh
Nice, thanks @mattezell yeah the sed stuff is almost always unreliable, wish the llms would leave it alone. Was going to update this PR with the latest tweaks, but it's mostly a throwaway nudge to get official support. We'll see what happens...
Thanks to you @tyeth - after reading your breakdown, I just couldn't resist the itch to validate the approach 🤣 Just had to see it working...
- added a commit that references this issue
on Mar 17, 2026 Closing this for now. As part of #873 we'll be moving off the local k3s implementation and onto more native solution on podman/docker/microvm. This should improve this class of bug.
- removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Oct 1, 2026
Description
OpenShell gateway with
--gpufails to make GPUs available to sandboxes when running on WSL2. Thenvidia-device-pluginDaemonSet either never schedules (0/0 replicas) or crashes withFailed to initialize NVML: Not Supported.Environment
Steps to Reproduce
openshell gateway start --gpunvidia-device-pluginDaemonSet has 0/0 desired replicasfeature.node.kubernetes.io/pci-10de.present=trueFailed to initialize NVML: Not SupportedRoot Cause
Three cascading issues on WSL2:
1. NFD cannot detect NVIDIA PCI device
WSL2 does not expose PCI topology to the guest kernel. NFD only sees
pci-1414.present(Microsoft Hyper-V), neverpci-10de.present(NVIDIA). The device plugin DaemonSet's node affinity is never satisfied.2. nvidia runtime
automode uses legacy injectionThe nvidia container runtime defaults to
mode = "auto"which selects legacy injection vianvidia-container-cli. On WSL2, this path does not properly injectlibdxcore.sointo pods — the critical library that bridges Linux NVML to the Windows DirectX GPU Kernel via/dev/dxg.3.
nvidia-ctk cdi generatemisseslibdxcore.soWhen generating the CDI spec,
nvidia-ctkcorrectly auto-detects WSL mode and selects/dev/dxg, but logs"Could not locate libdxcore.so"despite the library being present at/usr/lib/x86_64-linux-gnu/libdxcore.so. This is an upstream nvidia-container-toolkit bug (filed separately).Verified Fix
All three changes are required:
After these changes:
nvidia-device-plugin: 1/1 Runninggpu-feature-discovery: 1/1 Runningnvidia.com/gpu: 1nvidia-smiworks inside pods, RTX 5070 fully accessibleSuggested Permanent Fix
The
cluster-entrypoint.shshould detect WSL2 and apply these automatically:/dev/dxgexists or kernel version containsWSL2)nvidia-ctk cdi generate(already auto-detects WSL mode)libdxcore.somode = "cdi"in/etc/nvidia-container-runtime/config.tomlpci-10de.present=trueThis aligns with #398 (migrating to CDI for GPU injection) — WSL2 is a concrete platform where the legacy runtime stack is broken and CDI is the only viable path.
Agent Investigation
Diagnosed using
openshell doctorcommands:openshell doctor check— passed (Docker OK)openshell doctor exec -- kubectl get daemonset -n nvidia-device-plugin— 0/0 desiredopenshell doctor exec -- kubectl -n nvidia-device-plugin logs <pod>—NVML: Not Supportedopenshell doctor exec -- nvidia-ctk cdi generate— auto-detected WSL, warned about missing libdxcore.soopenshell doctor exec -- nvidia-container-cli info— successfully detected GPU at gateway level