Repository navigation
GetInferenceRoute intermittently returns NOT_FOUND (~51% of calls), aborting dcode -n non-interactive turns #2915
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Aug 25, 2026 📋 triage-agent
Classification: needs-information
The report shows a material failure pattern, but the supplied gateway entries do not establish that the successful and
NOT_FOUNDRPCs queried the same route in the same authenticated workspace.GetInferenceRoutereturnsNOT_FOUNDwhen the requested route is not configured in the resolved workspace. Also, HTTP status 200 with gRPC status 5 is expected gRPC response framing; the gRPC status is the operation result.Please provide, for one adjacent success/failure pair:
- Redacted
GetInferenceRouteRequestvalues for both calls (workspaceandroute_name) and the full gRPC status/message. - The authenticated principal or sandbox identity for both requests, or gateway diagnostics that establish it.
- The output of
openshell inference getimmediately after a failure, using the same gateway credentials and workspace as the NemoClaw integration. - The route/provider setup commands and whether any concurrent process can set, delete, or switch the inference route or workspace.
- A retest on current OpenShell v0.0.111, with the NemoClaw/dcode versions used.
This will distinguish a missing route or workspace/identity mismatch from an intermittent route-store defect.
- Redacted
- addedstate:needs-infoAssessment needs specific evidence or reproduction detailsAssessment needs specific evidence or reproduction detailsand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Aug 25, 2026 Thanks for the detailed ask. I re-pulled every claim in the original report directly from
our still-intact gateway log (it hasn't rotated since before this incident) rather than
re-asserting from memory, and added what's happened since. Splitting this into
re-verified original evidence (fact, re-derived independently just now) and new
findings since filing (also fact) from our working hypothesis connecting them
(clearly marked, not proven).Re-verified original evidence (2026-08-24)
Recomputed straight from
openshell-gateway.log, which still covers this window
unmodified:- 104 total
GetInferenceRoutecalls, 53 NOT_FOUND = 51.0%, spanning 02:37-23:53 UTC —
matches the original report exactly. - All 7 example pairs cited in the original report are still present, byte-for-byte,
including the cited 23:53:04 pair:and the other six at 22:42:24, 22:43:08, 22:45:00, 22:45:10, 23:22:28, 23:51:50 — all2026-08-24T23:53:04.440400Z GetInferenceRoute request_id=999a794f-24f0-468d-8665-81e5c0564a15 http=200 (OK) 2026-08-24T23:53:04.441564Z GetInferenceRoute request_id=d302fc56-930c-4eaa-a392-7d4ea33b1fca http=200 grpc.status_code=5 ERROR
independently re-confirmed present and unchanged. - The pattern was already fully deterministic then: of 52 pairs, 51 show
first-call-succeeds / second-call-NOT_FOUND (~0.9-1.7ms apart, distinct request_ids each
time), 1 pair both-fail, 0 pairs both-succeed. - Survived two
Starting OpenShell serverrestarts logged at 02:22:08 and 02:35:05, as the
original report noted.
New findings since filing
-
The identical pattern has continued, unchanged, ever since — including through a
third gateway restart (17:58:36) and a complete provider/model swap. Today's sample:
208 total calls, 108 NOT_FOUND (51.9%), 103/104 pairs still first-succeeds/second-fails,
~1.0-1.3ms apart. Same shape, different day, different models, different provider. -
But zero of those recent NOT_FOUND events correlate with an actual turn abort. Our
own orchestration log shows everyUnexpected error/managed non-interactive error
we've had was confined to a separate, unrelated 2026-08-25 14:39-18:00 UTC window (a
self-inflicted retry storm against our upstream provider, zeroGetInferenceRoutecalls
even occurred in that window) — and zero turn aborts in the 200+ cycles since
2026-08-25 19:40 UTC, despite the paired-NOT_FOUND pattern firing repeatedly and
identically throughout. So whatever made Aug 24 turns abort, it doesn't reproduce here
even though the RPC-level symptom you'd use to detect it does. -
What was different on 2026-08-24: we had three sandboxes (
gjh-coder,
gjh-reviewer,gjh-submitter) on one gateway, independently onboarded with three
different models — exactly the three named in the original report. NemoClaw's own
onboarding CLI told us why that's meaningful, verbatim from our terminal scrollback:$ nemo-deepagents gjh-coder connect Warning: gateway inference route (nvidia-prod/nvidia/nemotron-3-super-120b-a12b) differs from the recorded route for sandbox 'gjh-coder' (nvidia-prod/nvidia/nemotron-3-ultra-550b-a55b). Aligning the gateway to nvidia-prod/nvidia/nemotron-3-ultra-550b-a55b...Warning: Onboarding 'gjh-reviewer' will re-point the one shared inference route on OpenShell gateway 'nemoclaw' to nvidia-prod/nvidia/nemotron-3-super-120b-a12b. Affected registered sandboxes: 'gjh-coder'... OpenShell currently exposes this route per gateway, not per sandbox.Since 2026-08-25 19:40 UTC we've had all three sandboxes share one model on one route —
no more onboarding-triggered realignment. That's also exactly when turn aborts stopped. -
Direct evidence of write contention on this store during onboarding: our gateway's
own trace log has 14sqlx::query: slow statement: execution time exceeded alert thresholdWARNs. Two, both timestamped during onboarding activity:store.get_by_name object_type=inference_route workspace=default object.name=inference.local elapsed=1.529886341s store.put_if object_type=sandbox object.id=... object.name=gjh-coder workspace=default elapsed=5.062471243sA 1.5s single-row indexed SELECT and a 5s+ resource_version-guarded UPDATE against the
same SQLite-backedobjectsstore, both during the exact activity type (multi-sandbox
onboarding) that produced the original report's 3-model reproduction. I checked whether
these 14 slow-statement events line up with the 108 NOT_FOUND timestamps directly: only
1 of 108 falls within 30 seconds of one. So this contention doesn't explain the raw
per-call NOT_FOUND rate — it's a separate, real phenomenon that plausibly explains
route mismatches during onboarding, not the deterministic per-call pairing itself.
Our working hypothesis (not proven — flagging clearly as interpretation)
Two independent things appear to be going on, and we conflated them in the original
report:- (a) A deterministic, apparently benign RPC pattern: every logical
GetInferenceRouteresolution fires two calls, one always succeeding and one always
NOT_FOUND, ~1ms apart. Reproducible on demand with plainopenshell inference get
(no special flags) — that command prints both an "Inference" and a "System inference"
section from one invocation, and shows "System inference: Not configured" for us without
ever surfacing an error. We think the second call in each pair is a distinct, and on our
setup legitimately-unconfigured, secondary route scope — not a flaky re-read of the same
route. This pattern has never once correlated with a turn abort in anything we've been
able to observe. - (b) The actual cause of Aug 24's aborts: three sandboxes independently onboarding
different models onto the single gateway-level route, with real (seconds-long, logged)
SQLite write contention on that shared row during the churn. We can't prove (a) and (b)
are unrelated — both surface at roughly the same "GetInferenceRoute gets called" moments
— but (a) alone, without (b)'s multi-model contention, produces zero aborts in 200+
cycles of continuous operation since.
If that split is right, the fix isn't in the NOT_FOUND pairing itself — it's in making
GetInferenceRouteresilient to reading a route that's mid-realignment during another
sandbox's onboarding, or genuinely scoping routes per-sandbox as the sandbox-level config
already implies.Your five points, directly
- Adjacent pair / redacted request values: see the re-verified 23:53:04 pair above.
I can't give you the literalGetInferenceRouteRequest.workspace/route_namefields —
our structured logging captures RPC method + status only, not request bodies, and
openshell -vvvdoesn't log payload-level detail either (checked directly, only TLS
handshake detail at that verbosity). - Authenticated principal: single shared identity for every call across all sandboxes
—openshell whoami→ Subjectopenshell-client, Providermtls. openshell inference getright after a failure:Version 13 = 13Inference: Workspace: default | Provider: compatible-endpoint | Model: nemotron-3-ultra-free | Version: 13 System inference: Not configuredSetInferenceRoutewrites logged against this one row total.- Setup commands / concurrency:
nemo-deepagents <sandbox> connect/onboard/
rebuildagainst our three sandboxes — nothing else touchesinference set/update/delete. But by the gateway's own design, onboarding one sandbox does
concurrently affect the shared route used by the others — see the warnings quoted
above. Our own orchestrator runs sandboxes sequentially under aflock, so
cross-sandbox concurrency on our side is only in play during interactive
onboarding/connect sessions, not our scheduled loop. - Retest on v0.0.111: not reachable from here —
nemoclaw update --checkreports
0.0.109 as the latest maintained version, bundling OpenShell 0.0.101. No upgrade path
found to 0.0.111 through the maintained installer.
Happy to pull anything else out of the gateway log — it's still intact back to before this
incident started.- 104 total
This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Sep 9, 2026 📋 triage-agent
Closing as obsolete after #3195. That merged change removed the managed inference control plane, route persistence,
GetInferenceRoute, and theinference.localdata path. The RPC and route records implicated by this report no longer exist; current inference workloads attach provider profiles and call provider-native endpoints.
User Story
As an operator running multiple long-lived NemoClaw sandboxes (LangChain Deep Agents Code /
dcode) against NVIDIA-hosted inference for an unattended, cron-driven agent loop, I needdcode -nnon-interactive turns to complete reliably so that a multi-step agentic task (several tool calls in one turn) isn't aborted mid-way by the gateway itself.Problem Statement
Roughly half of all
dcode -nnon-interactive turns that make more than a couple of tool calls abort withUnexpected error (correlation_id=...)insidedcode, and the sandbox'snemoclaw <name> exec ... -- dcode -n "..."wrapper reportsmanaged non-interactive error: error_class=unknown category=unknown retryable=false.Correlating the gateway's own request log (
openshell-gateway.log) against every failure'scorrelation_id, each failure lines up with the same internal pattern: the gateway issues aPOST /openshell.inference.v1.Inference/GetInferenceRoutecall that succeeds (status=200, no gRPC error), then immediately (within ~1ms) issues a secondGetInferenceRoutecall for the same request that returnsrpc.grpc.status_code=5(NOT_FOUND) despitehttp.response.status_code=200. This paired-call/second-fails pattern is the only anomaly in the log at each failure timestamp.Impact / Why This Matters
GetInferenceRoutecalls (51%) returned thisNOT_FOUND, spanning from 02:22 UTC (before any of my sessions started) through the most recent attempt, and surviving an OpenShell server restart that already happened in between (Starting OpenShell serverlogged twice: 02:22:08 and 02:35:05).nvidia/nemotron-3-ultra-550b-a55b,nvidia/nemotron-3-super-120b-a12b, andminimaxai/minimax-m3— ruling out a model-specific cause; the failure is below the model-selection layer.dcode -nturn from scratch (a fresh gateway session/correlation_id) — a same-turn retry is not possible sinceretryable=false. In an unattended/cron-driven setup this means ~50% of scheduled work cycles are silently wasted and must wait for the next cron tick, roughly doubling end-to-end latency for any multi-tool-call task.nemoclawCLI:nemoclaw <name> gateway restartis unsupported for thelangchain-deepagents-codeagent ("LangChain Deep Agents Code has no gateway runtime"), and restarting the shared OpenShell server did not clear the underlying issue earlier today.Acceptance Criteria
dcode -nnon-interactive turn with several sequential tool calls no longer aborts withUnexpected error/managed non-interactive error: error_class=unknowndue to aGetInferenceRouteNOT_FOUND.GetInferenceRoutecalls succeed consistently for a sandbox with a valid, active inference route (no spurious immediate re-call/NOT_FOUNDpair after a successful200/no-error response).dcodeclassifies asretryable=true, rather than aborting the whole turn.Reproduction Steps
nemoclaw onboard --non-interactive --yes --name <sandbox> --agent langchain-deepagents-code ...against an NVIDIA-hosted model (nvidia-prod/buildprovider).nemoclaw <sandbox> exec --workdir /sandbox/<project> --timeout 1500 -- dcode -n "<a task prompt requiring the model to read several files via read_file/ls before writing anything>"several times in a row (a handful of runs is enough to reproduce; roughly 1 in 2 fails).dcode's own output: several tool calls succeed, thenUnexpected error (correlation_id=<id>)and the process exits non-zero.~/.local/state/nemoclaw/openshell-docker-gateway/openshell-gateway.logaround that timestamp and look for aGetInferenceRouterequest immediately followed by anotherGetInferenceRouterequest withrpc.grpc.status_code=5.Environment
langchain-deepagents-code(dcodev0.1.34) sandboxes, inference providerbuild(NVIDIA-hosted), modelsnvidia/nemotron-3-ultra-550b-a55b,nvidia/nemotron-3-super-120b-a12b,minimaxai/minimax-m3— same failure on all three.Logs
Matching pair from
openshell-gateway.logat the same time as one such failure (request IDs and sandbox IDs are random UUIDs, not credentials):Six more occurrences of the identical paired-call pattern were observed at 22:42:24, 22:43:08, 22:45:00, 22:45:10, 23:22:28, and 23:51:50 UTC in the same session, each within ~1-2ms of a
dcodeturn aborting.