Skip to content

feat(observability): OpenTelemetry export surface for gateway operators #2507

Description

@krishicks

Problem Statement

OpenShell gives gateway operators three observability surfaces today, and none of them answer "what is my gateway doing right now, and why is it slow or failing?"

  • /metrics (Prometheus) exposes gRPC/HTTP request counters and duration histograms, a readiness gauge, and gateway-interceptor counters. It is scrape-only, disabled by default (--metrics-port defaults to 0), and has no Helm scrape wiring.
  • OCSF logs (shorthand always on, full JSONL opt-in) describe sandbox security events for SIEM consumption. They are not gateway operational telemetry and carry no timing or causality.
  • Anonymous product telemetry (openshell-core::telemetry) forwards aggregate usage to NVIDIA. It is deliberately coarse and is not readable by the operator running the gateway.

What is missing is the connective tissue operators expect from any modern control plane: distributed traces, an OTLP push path for environments that do not scrape, and a turnkey collector wiring so this works without hand-assembling a stack.

Concretely, an operator investigating a slow CreateSandbox today can see that the request took 4.2s from the request-duration histogram, and can find its log lines via the request ID from #932 — but cannot see where those 4.2 seconds went across compute-driver call, policy evaluation, supervisor session establishment, and readiness wait. That decomposition is exactly what a trace provides and no current surface does.

This issue is the tracking parent for that work. It is a sub-issue of #1055, which remains the roadmap-level umbrella covering OCSF, log export, product telemetry, and dashboards.

Audience: the operator running an OpenShell gateway. Tracing for users of the SDKs who create sandboxes (#1818) is explicitly out of scope here.

Proposed Design

Deliver an OTel export surface for the gateway process, decomposed into sub-issues. PR #1270 built a working version of most of this and was closed unmerged on 2026-05-27; it is a valuable reference, but the implementation should be re-derived rather than rebased (see Agent Investigation for why).

Proposed sub-issues

  1. Gateway OTLP trace export. Append a tracing-opentelemetry layer to the existing subscriber in tracing_bus.rs, so the current tower_http::trace::TraceLayer per-request span becomes the OTLP root without rewriting handlers. Resolve configuration from standard OTEL_* environment variables with a CLI/TOML fallback, honor OTEL_TRACES_SAMPLER, and flush the batch span processor on shutdown. Default off.

  2. Span coverage for gateway operations. Instrument the spans that make a trace useful rather than merely present: sandbox lifecycle transitions, compute-driver calls, policy evaluation, supervisor session establishment and relay claim, and credential-broker actions. Decide per operation whether it is a span, an event on a parent span, or neither.

  3. Correlation identifiers. Carry stable fields across spans and align them with OCSF event correlation IDs so an operator can pivot from a trace to the security record for the same activity, and reuse the request ID from feat(server): add request-ID middleware for request correlation #932 rather than inventing a parallel identifier. This subsumes the emission half of Add OpenTelemetry trace correlation across gateway activity #1758.

  4. OTLP metrics export. Offer OTLP push for the existing metrics families alongside the current Prometheus scrape endpoint, for operators whose collectors do not scrape. Per the discussion on feat(server): metrics instrumentation #909, scrape remains the default and push is opt-in.

  5. Kubernetes monitoring surface. Helm ServiceMonitor (gated, off by default) targeting the existing named metrics port, OTEL_* projection into the StatefulSet, and values documentation. Also decide whether --metrics-port should default to enabled when the chart wires up scraping.

  6. Local development stack. A mise task installing a collector plus a trace backend and Grafana into the k3d dev cluster, so contributors can validate the pipeline end to end. Mirrors the existing Keycloak dev add-on pattern.

  7. Documentation. An operator-facing page under docs/observability/ covering enablement, configuration surface, what spans exist, and a reference collector setup — plus an update to architecture/gateway.md.

Sub-issues are expected to land independently; the trace export path (1–3) is the critical path and the rest can follow.

Open design questions

  • Config surface. OTEL_* env vars are the ecosystem convention, but OpenShell has a TOML gateway config (RFC 0003) that should probably own this. Precedence between the two needs a decision, and docs/reference/gateway-config.mdx needs updating either way.
  • Dependency cost. The OTel Rust SDK plus OTLP exporter is a non-trivial addition to the gateway's dependency graph and SBOM. Feature-gating at compile time is worth evaluating, especially alongside Conditional compilation of compute drivers #1943.
  • Sandbox boundary. Whether supervisor-side spans join the gateway trace is deliberately left out of scope. It interacts with Push gateway-owned desired state to supervisors #1731's supervisor-session restructuring and with sandbox egress policy, and deserves its own design.

Alternatives Considered

  • Rebase PR feat(observability): add gateway OTLP traces and initial Kube monitoring surface #1270 as-is. Fastest path, but its dependency pins were chosen for a workspace state that no longer exists, and it predates the gateway TOML config and the interceptor and middleware subsystems. The design is sound; the diff is stale.
  • Metrics only, finish feat(server): metrics instrumentation #909's catalog. Cheaper and complementary, and should happen regardless — but aggregate counters cannot decompose a single slow request across subsystems, which is the specific gap here.
  • Rely on OCSF JSONL for operational analysis. OCSF describes sandbox security activity for SIEM consumption. It is the wrong schema and the wrong granularity for gateway latency and failure analysis.
  • Defer to a service mesh. Mesh-level traces capture the network hop but not policy evaluation, driver calls, or session establishment inside the gateway process, and would impose a mesh dependency on all deployments.

Agent Investigation

Verified against main at 0d5e5c5.

No OpenTelemetry exists in the repository. grep -rni "opentelemetry|otlp" across crates/, python/, deploy/, and docs/ returns a single hit, in deploy/sbom/resolve_licenses.py, unrelated to product code.

What is present:

  • metrics facade + metrics-exporter-prometheus, /metrics route at crates/openshell-server/src/http.rs:172, builder init at crates/openshell-server/src/lib.rs:472.
  • Emitted families are limited to openshell_server_grpc_requests_total, openshell_server_grpc_request_duration_seconds, openshell_server_http_requests_total, openshell_server_http_request_duration_seconds (multiplex.rs), a readiness gauge (readiness.rs), and interceptor fail-open/fail-closed/latency (openshell-gateway-interceptors/src/runtime.rs). The broader catalog proposed in feat(server): metrics instrumentation #909 is unimplemented.
  • --metrics-port / OPENSHELL_METRICS_PORT defaults to 0, meaning the dedicated metrics listener is off unless configured (cli.rs:68).
  • Request-ID correlation from feat(server): add request-ID middleware for request correlation #932 is in place; that issue framed itself as "the minimum unit of correlation before adopting full OpenTelemetry."

Why PR #1270 should be re-derived rather than rebased: it pinned opentelemetry 0.29 and tracing-opentelemetry 0.30, described in the PR body as the latest set compatible with the workspace's tonic 0.12 + prost 0.13. The workspace is now on tonic 0.14 / tonic-prost 0.14 / prost 0.14, so that pin rationale no longer holds and the version selection has to be redone. The PR was also authored before the gateway TOML config (#1317), the gateway interceptors (#2005), and the supervisor middleware crates landed — all of which are plausible span sources. It was closed with no review comments explaining the decision, so the reason for abandonment is not recoverable from the PR itself and may be worth confirming with a maintainer before investing.

Related issues:

Deliberately adjacent, not children:

Checklist

  • I've reviewed existing issues and the architecture docs
  • This is a design proposal, not a "please build this" request

Activity

  1. added
    topic:observabilityLogging, metrics, and observability work
    area:gatewayGateway server and control-plane work
    area:clusterRelated to running OpenShell on k3s/docker
    on Jul 27, 2026
  2. krishicks commented on Jul 27, 2026

    @krishicks
    CollaboratorAuthor

    Additional sub-issue: compute driver and gateway interceptor spans

    Adding an eighth proposed sub-issue, surfaced while scoping the sandbox-side companion #2508.

    8. Compute driver and gateway interceptor spans, with trace context propagation.

    Both are gateway-side surfaces — the gateway is the client in each case — so they belong here rather than in #2508.

    Compute drivers. This is the missing piece of the CreateSandbox decomposition in the problem statement above. Without driver spans, "sandbox creation took 4.2s" bottoms out at "the driver took 3s," with no visibility into whether the time went to image pull, scheduling, or readiness wait.

    Note this is a moving target: drivers are expected to move from in-process backends to external gRPC services over proto/compute_driver.proto, while remaining in-tree. Today only the VM driver is a separate process (crates/openshell-server/src/compute/vm.rs); Docker, Podman, and Kubernetes are in-process. Once that migration happens, every driver call becomes a real process boundary and compute_driver.proto becomes a trace-context propagation point.

    That has two consequences for how this sub-issue should be built:

    • Do not design driver instrumentation as in-process spans that later need rewriting. Treat the driver contract as a propagation boundary from the start, even for backends that are currently in-process, so the same instrumentation survives the migration.
    • Trace context propagation over compute_driver.proto should be settled as part of that contract's evolution rather than bolted on afterward. Worth coordinating with whoever owns the external-driver migration before implementing.

    Gateway interceptors. Evaluate on proto/gateway_interceptor.proto already calls an out-of-process, operator-run service, so this is a propagation boundary today. The existing openshell_gateway_interceptor_fail_closed_total, fail_open_total, and latency_seconds metrics (crates/openshell-gateway-interceptors/src/runtime.rs) are evidence the observability need was already felt here, and they give a ready-made baseline to validate span timings against.

    Both cases share a property that makes them high value: the operator controls the service on the other side, so propagating context yields a genuine cross-service trace rather than a span that dead-ends at the boundary. This is the same argument that makes supervisor middleware the strongest case in #2508.

  3. elezar commented on Jul 28, 2026

    @elezar
    Member

    Concrete motivator from #2516 / #2498 follow-up:

    The rootless Podman E2E lane on main timed out after #2498 raised the production Podman graceful stop timeout to 45s. The tests had already passed; the remaining time was spent in gateway/sandbox teardown. #2516 fixes CI by overriding the Podman E2E harness timeout to 15s while preserving the production default.

    This is a good example of why compute-driver lifecycle spans belong in this observability work. For this class of issue, operators need to see teardown decomposed by phase, not just the outer sandbox delete duration or warning logs. Useful span attributes/events would include:

    • driver.name
    • sandbox.id
    • operation = stop|remove|cleanup
    • timeout_secs
    • elapsed_ms
    • outcome = graceful|forced|not_found|error
    • backend-specific identifiers such as container/pod/vm id

    For future external drivers, this also reinforces the need for trace context propagation across compute_driver.proto: the gateway can record the outer driver RPC span, while the driver records child spans for backend-specific phases such as Podman stop, Kubernetes pod deletion, or VM shutdown.

  4. 21 remaining items

  5. rhuss commented on Aug 6, 2026

    @rhuss
    Contributor

    Persona workflows from the observability RFC work (PRs #2628/#2629, now closed in favor of issues). These workflows motivate the driver instrumentation and deployment configuration tracked here:

    Platform operator: debugging slow sandbox creation

    An operator receives an alert that sandbox creation latency exceeds the SLO. They open Jaeger and search for sandbox.create spans in the last hour. They find one taking 15 seconds. They drill into the child spans, expecting to see which compute driver operation was slow, but the trace ends at the gateway request span. The K8s driver's provision_pod call, the Docker driver's create_container, the Podman driver's equivalent are all invisible because none of them have #[tracing::instrument] annotations.

    How this issue enables the workflow: The in-process driver spans appear as children of the gateway request span. The operator sees that provision_pod took 14 of the 15 seconds, drills into the K8s API calls, and identifies that the node was under memory pressure. No new infrastructure is needed because these drivers already inherit the gateway's tracing subscriber.

    Platform integrator: wiring OpenShell into an existing monitoring stack

    A platform integrator is deploying OpenShell on a managed K8s cluster that already runs Prometheus, Grafana, and Tempo. They need OpenShell to emit traces to their Tempo endpoint and expose Prometheus metrics for their existing dashboards and alerts. Today, the Helm chart has no OTLP configuration values, no ServiceMonitor template, and the TOML config reference does not document the OTLP section.

    How this issue enables the workflow: The integrator sets otlp.enabled=true and otlp.endpoint=http://tempo:4317 in the Helm values. They enable the ServiceMonitor and Prometheus starts scraping the gateway's /metrics endpoint. OpenShell traces appear in Tempo alongside their other services. The published docs at docs/reference/gateway-config.mdx document the configuration surface.

  6. github-actions commented on Aug 23, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:clusterRelated to running OpenShell on k3s/dockerarea:gatewayGateway server and control-plane workstate:staleInactive item at risk of automatic closure.topic:observabilityLogging, metrics, and observability work

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions