Skip to content

GetInferenceRoute intermittently returns NOT_FOUND (~51% of calls), aborting dcode -n non-interactive turns #2915

Description

@reqsgjh

User Story

As an operator running multiple long-lived NemoClaw sandboxes (LangChain Deep Agents Code / dcode) against NVIDIA-hosted inference for an unattended, cron-driven agent loop, I need dcode -n non-interactive turns to complete reliably so that a multi-step agentic task (several tool calls in one turn) isn't aborted mid-way by the gateway itself.

Problem Statement

Roughly half of all dcode -n non-interactive turns that make more than a couple of tool calls abort with Unexpected error (correlation_id=...) inside dcode, and the sandbox's nemoclaw <name> exec ... -- dcode -n "..." wrapper reports managed non-interactive error: error_class=unknown category=unknown retryable=false.

Correlating the gateway's own request log (openshell-gateway.log) against every failure's correlation_id, each failure lines up with the same internal pattern: the gateway issues a POST /openshell.inference.v1.Inference/GetInferenceRoute call that succeeds (status=200, no gRPC error), then immediately (within ~1ms) issues a second GetInferenceRoute call for the same request that returns rpc.grpc.status_code=5 (NOT_FOUND) despite http.response.status_code=200. This paired-call/second-fails pattern is the only anomaly in the log at each failure timestamp.

Impact / Why This Matters

  • Across today's session, 53 of 104 GetInferenceRoute calls (51%) returned this NOT_FOUND, spanning from 02:22 UTC (before any of my sessions started) through the most recent attempt, and surviving an OpenShell server restart that already happened in between (Starting OpenShell server logged twice: 02:22:08 and 02:35:05).
  • Reproduced identically across three different backing models on three separate sandboxes — nvidia/nemotron-3-ultra-550b-a55b, nvidia/nemotron-3-super-120b-a12b, and minimaxai/minimax-m3 — ruling out a model-specific cause; the failure is below the model-selection layer.
  • Practically, users must retry the entire dcode -n turn from scratch (a fresh gateway session/correlation_id) — a same-turn retry is not possible since retryable=false. In an unattended/cron-driven setup this means ~50% of scheduled work cycles are silently wasted and must wait for the next cron tick, roughly doubling end-to-end latency for any multi-tool-call task.
  • No workaround exists inside the sandbox or via nemoclaw CLI: nemoclaw <name> gateway restart is unsupported for the langchain-deepagents-code agent ("LangChain Deep Agents Code has no gateway runtime"), and restarting the shared OpenShell server did not clear the underlying issue earlier today.

Acceptance Criteria

  • A dcode -n non-interactive turn with several sequential tool calls no longer aborts with Unexpected error / managed non-interactive error: error_class=unknown due to a GetInferenceRoute NOT_FOUND.
  • GetInferenceRoute calls succeed consistently for a sandbox with a valid, active inference route (no spurious immediate re-call/NOT_FOUND pair after a successful 200/no-error response).
  • If a route genuinely becomes stale/invalid mid-session, the gateway either recovers automatically or returns an error that dcode classifies as retryable=true, rather than aborting the whole turn.

Reproduction Steps

  1. nemoclaw onboard --non-interactive --yes --name <sandbox> --agent langchain-deepagents-code ... against an NVIDIA-hosted model (nvidia-prod / build provider).
  2. Run nemoclaw <sandbox> exec --workdir /sandbox/<project> --timeout 1500 -- dcode -n "<a task prompt requiring the model to read several files via read_file/ls before writing anything>" several times in a row (a handful of runs is enough to reproduce; roughly 1 in 2 fails).
  3. Observe dcode's own output: several tool calls succeed, then Unexpected error (correlation_id=<id>) and the process exits non-zero.
  4. Tail ~/.local/state/nemoclaw/openshell-docker-gateway/openshell-gateway.log around that timestamp and look for a GetInferenceRoute request immediately followed by another GetInferenceRoute request with rpc.grpc.status_code=5.

Environment

  • NemoClaw: v0.0.109
  • OpenShell: 0.0.101 (docker driver)
  • OS: Ubuntu 24.04.4 LTS
  • Docker: 29.1.3 (build 29.1.3-0ubuntu3~24.04.2)
  • Runtime/integration: NemoClaw-managed langchain-deepagents-code (dcode v0.1.34) sandboxes, inference provider build (NVIDIA-hosted), models nvidia/nemotron-3-ultra-550b-a55b, nvidia/nemotron-3-super-120b-a12b, minimaxai/minimax-m3 — same failure on all three.

Logs

managed non-interactive error: error_class=unknown category=unknown retryable=false correlation_id=01a03631-51a5-7053-a922-5557a8a6f9d0
...
Unexpected error (correlation_id=01a03631-51a5-7053-a922-5557a8a6f9d0)

Matching pair from openshell-gateway.log at the same time as one such failure (request IDs and sandbox IDs are random UUIDs, not credentials):

2026-08-24T23:53:04.440400Z INFO request{method=POST path="/openshell.inference.v1.Inference/GetInferenceRoute" request_id="999a794f-24f0-468d-8665-81e5c0564a15" ... http.response.status_code=200}: response status=200 latency_ms=0
2026-08-24T23:53:04.441564Z INFO request{method=POST path="/openshell.inference.v1.Inference/GetInferenceRoute" request_id="d302fc56-930c-4eaa-a392-7d4ea33b1fca" ... http.response.status_code=200 rpc.grpc.status_code=5 otel.status_code="ERROR"}: response status=200 latency_ms=0

Six more occurrences of the identical paired-call pattern were observed at 22:42:24, 22:43:08, 22:45:00, 22:45:10, 23:22:28, and 23:51:50 UTC in the same session, each within ~1-2ms of a dcode turn aborting.

Activity

  1. elezar commented on Aug 25, 2026

    @elezar
    Member

    📋 triage-agent

    Classification: needs-information

    The report shows a material failure pattern, but the supplied gateway entries do not establish that the successful and NOT_FOUND RPCs queried the same route in the same authenticated workspace.

    GetInferenceRoute returns NOT_FOUND when the requested route is not configured in the resolved workspace. Also, HTTP status 200 with gRPC status 5 is expected gRPC response framing; the gRPC status is the operation result.

    Please provide, for one adjacent success/failure pair:

    1. Redacted GetInferenceRouteRequest values for both calls (workspace and route_name) and the full gRPC status/message.
    2. The authenticated principal or sandbox identity for both requests, or gateway diagnostics that establish it.
    3. The output of openshell inference get immediately after a failure, using the same gateway credentials and workspace as the NemoClaw integration.
    4. The route/provider setup commands and whether any concurrent process can set, delete, or switch the inference route or workspace.
    5. A retest on current OpenShell v0.0.111, with the NemoClaw/dcode versions used.

    This will distinguish a missing route or workspace/identity mismatch from an intermittent route-store defect.

  2. added
    state:needs-infoAssessment needs specific evidence or reproduction details
    and removed
    state:triage-neededOpened without agent diagnostics and needs triage
    on Aug 25, 2026
  3. reqsgjh commented on Aug 26, 2026

    @reqsgjh
    Author

    Thanks for the detailed ask. I re-pulled every claim in the original report directly from
    our still-intact gateway log (it hasn't rotated since before this incident) rather than
    re-asserting from memory, and added what's happened since. Splitting this into
    re-verified original evidence (fact, re-derived independently just now) and new
    findings since filing
    (also fact) from our working hypothesis connecting them
    (clearly marked, not proven).

    Re-verified original evidence (2026-08-24)

    Recomputed straight from openshell-gateway.log, which still covers this window
    unmodified:

    • 104 total GetInferenceRoute calls, 53 NOT_FOUND = 51.0%, spanning 02:37-23:53 UTC —
      matches the original report exactly.
    • All 7 example pairs cited in the original report are still present, byte-for-byte,
      including the cited 23:53:04 pair:
      2026-08-24T23:53:04.440400Z  GetInferenceRoute  request_id=999a794f-24f0-468d-8665-81e5c0564a15  http=200 (OK)
      2026-08-24T23:53:04.441564Z  GetInferenceRoute  request_id=d302fc56-930c-4eaa-a392-7d4ea33b1fca  http=200 grpc.status_code=5 ERROR
      
      and the other six at 22:42:24, 22:43:08, 22:45:00, 22:45:10, 23:22:28, 23:51:50 — all
      independently re-confirmed present and unchanged.
    • The pattern was already fully deterministic then: of 52 pairs, 51 show
      first-call-succeeds / second-call-NOT_FOUND (~0.9-1.7ms apart, distinct request_ids each
      time), 1 pair both-fail, 0 pairs both-succeed.
    • Survived two Starting OpenShell server restarts logged at 02:22:08 and 02:35:05, as the
      original report noted.

    New findings since filing

    1. The identical pattern has continued, unchanged, ever since — including through a
      third gateway restart (17:58:36) and a complete provider/model swap. Today's sample:
      208 total calls, 108 NOT_FOUND (51.9%), 103/104 pairs still first-succeeds/second-fails,
      ~1.0-1.3ms apart. Same shape, different day, different models, different provider.

    2. But zero of those recent NOT_FOUND events correlate with an actual turn abort. Our
      own orchestration log shows every Unexpected error/managed non-interactive error
      we've had was confined to a separate, unrelated 2026-08-25 14:39-18:00 UTC window (a
      self-inflicted retry storm against our upstream provider, zero GetInferenceRoute calls
      even occurred in that window) — and zero turn aborts in the 200+ cycles since
      2026-08-25 19:40 UTC, despite the paired-NOT_FOUND pattern firing repeatedly and
      identically throughout. So whatever made Aug 24 turns abort, it doesn't reproduce here
      even though the RPC-level symptom you'd use to detect it does.

    3. What was different on 2026-08-24: we had three sandboxes (gjh-coder,
      gjh-reviewer, gjh-submitter) on one gateway, independently onboarded with three
      different models — exactly the three named in the original report. NemoClaw's own
      onboarding CLI told us why that's meaningful, verbatim from our terminal scrollback:

      $ nemo-deepagents gjh-coder connect
        Warning: gateway inference route (nvidia-prod/nvidia/nemotron-3-super-120b-a12b) differs
        from the recorded route for sandbox 'gjh-coder' (nvidia-prod/nvidia/nemotron-3-ultra-550b-a55b).
        Aligning the gateway to nvidia-prod/nvidia/nemotron-3-ultra-550b-a55b...
      
        Warning: Onboarding 'gjh-reviewer' will re-point the one shared inference route on
        OpenShell gateway 'nemoclaw' to nvidia-prod/nvidia/nemotron-3-super-120b-a12b. Affected
        registered sandboxes: 'gjh-coder'... OpenShell currently exposes this route per
        gateway, not per sandbox.
      

      Since 2026-08-25 19:40 UTC we've had all three sandboxes share one model on one route —
      no more onboarding-triggered realignment. That's also exactly when turn aborts stopped.

    4. Direct evidence of write contention on this store during onboarding: our gateway's
      own trace log has 14 sqlx::query: slow statement: execution time exceeded alert threshold WARNs. Two, both timestamped during onboarding activity:

      store.get_by_name object_type=inference_route workspace=default object.name=inference.local
        elapsed=1.529886341s
      store.put_if object_type=sandbox object.id=... object.name=gjh-coder workspace=default
        elapsed=5.062471243s
      

      A 1.5s single-row indexed SELECT and a 5s+ resource_version-guarded UPDATE against the
      same SQLite-backed objects store, both during the exact activity type (multi-sandbox
      onboarding) that produced the original report's 3-model reproduction. I checked whether
      these 14 slow-statement events line up with the 108 NOT_FOUND timestamps directly: only
      1 of 108 falls within 30 seconds of one. So this contention doesn't explain the raw
      per-call NOT_FOUND rate — it's a separate, real phenomenon that plausibly explains
      route mismatches during onboarding, not the deterministic per-call pairing itself.

    Our working hypothesis (not proven — flagging clearly as interpretation)

    Two independent things appear to be going on, and we conflated them in the original
    report:

    • (a) A deterministic, apparently benign RPC pattern: every logical
      GetInferenceRoute resolution fires two calls, one always succeeding and one always
      NOT_FOUND, ~1ms apart. Reproducible on demand with plain openshell inference get
      (no special flags) — that command prints both an "Inference" and a "System inference"
      section from one invocation, and shows "System inference: Not configured" for us without
      ever surfacing an error. We think the second call in each pair is a distinct, and on our
      setup legitimately-unconfigured, secondary route scope — not a flaky re-read of the same
      route. This pattern has never once correlated with a turn abort in anything we've been
      able to observe.
    • (b) The actual cause of Aug 24's aborts: three sandboxes independently onboarding
      different models onto the single gateway-level route, with real (seconds-long, logged)
      SQLite write contention on that shared row during the churn. We can't prove (a) and (b)
      are unrelated — both surface at roughly the same "GetInferenceRoute gets called" moments
      — but (a) alone, without (b)'s multi-model contention, produces zero aborts in 200+
      cycles of continuous operation since.

    If that split is right, the fix isn't in the NOT_FOUND pairing itself — it's in making
    GetInferenceRoute resilient to reading a route that's mid-realignment during another
    sandbox's onboarding, or genuinely scoping routes per-sandbox as the sandbox-level config
    already implies.

    Your five points, directly

    1. Adjacent pair / redacted request values: see the re-verified 23:53:04 pair above.
      I can't give you the literal GetInferenceRouteRequest.workspace/route_name fields —
      our structured logging captures RPC method + status only, not request bodies, and
      openshell -vvv doesn't log payload-level detail either (checked directly, only TLS
      handshake detail at that verbosity).
    2. Authenticated principal: single shared identity for every call across all sandboxes
      — openshell whoami → Subject openshell-client, Provider mtls.
    3. openshell inference get right after a failure:
      Inference: Workspace: default | Provider: compatible-endpoint | Model: nemotron-3-ultra-free | Version: 13
      System inference: Not configured
      
      Version 13 = 13 SetInferenceRoute writes logged against this one row total.
    4. Setup commands / concurrency: nemo-deepagents <sandbox> connect / onboard /
      rebuild against our three sandboxes — nothing else touches inference set/update/delete. But by the gateway's own design, onboarding one sandbox does
      concurrently affect the shared route used by the others — see the warnings quoted
      above. Our own orchestrator runs sandboxes sequentially under a flock, so
      cross-sandbox concurrency on our side is only in play during interactive
      onboarding/connect sessions, not our scheduled loop.
    5. Retest on v0.0.111: not reachable from here — nemoclaw update --check reports
      0.0.109 as the latest maintained version, bundling OpenShell 0.0.101. No upgrade path
      found to 0.0.111 through the maintained installer.

    Happy to pull anything else out of the gateway log — it's still intact back to before this
    incident started.

  4. github-actions commented on Sep 9, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

  5. johntmyers commented on Sep 17, 2026

    @johntmyers
    Collaborator

    📋 triage-agent

    Closing as obsolete after #3195. That merged change removed the managed inference control plane, route persistence, GetInferenceRoute, and the inference.local data path. The RPC and route records implicated by this report no longer exist; current inference workloads attach provider profiles and call provider-native endpoints.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:needs-infoAssessment needs specific evidence or reproduction detailsstate:staleInactive item at risk of automatic closure.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions