Repository navigation
Kubernetes Operator #1719
Description
Activity
- changed the title
[-]feat: explore OpenShell Kubernetes Operator design[/-][+]Kubernetes Operator[/+]on Jun 3, 2026 We operate an OpenShell Kubernetes Operator. Sharing feedback based on that experience.
Personas
Before answering the questions — a framing that shaped most of our design decisions. The issue treats "platform teams" and "users" as two groups. In practice we found three distinct personas with different trust levels:
Platform Admin (cluster operator): "Run OpenShell reliably, securely, at scale."
- Deploys the operator and its dependencies
- Manages CRDs, webhooks, cluster-level defaults (supervisor delivery method, image pull policies)
- Scopes Tenant Admins via namespace-level RBAC — but does not decide which providers or policies exist within a tenant's domain
Tenant Admin (DevOps / team lead): "Configure the sandbox environment for my team."
- Authors security policies (filesystem, network, process restrictions)
- Manages the credential plane (which API keys are available, how they're scoped)
- Sets infrastructure parameters (images, RuntimeClass, resource limits)
Sandbox Consumer (end user / developer / service): "Give me an isolated environment."
- Does not care about infrastructure, security policy, or credentials
- Two sub-cases worth distinguishing:
- A service requesting a sandbox (knows OpenShell, knows how to create a CR — e.g. a chatbot backend spinning up an agent sandbox for each session)
- An end user of that service (doesn't know OpenShell exists — e.g. someone chatting with the bot)
- In K8s API terms, the "Sandbox Consumer" is always the service or human creating the CR — the end user behind it is invisible to the operator
These personas should be answered for differently. A "user-facing Sandbox resource" means different things to the Sandbox Consumer (who just wants
kubectl apply -f my-sandbox.yamlwith a policy reference) versus the Tenant Admin (who defines what that policy allows).
Questions
Would you use an operator mainly to deploy OpenShell, to manage sandboxes, or both?
Both. A deployment-only operator becomes insufficient quickly — platform teams appreciate a declarative sandbox lifecycle from CI pipelines, GitOps tools, and custom controllers. Shipping both in a single operator with a clear boundary feels like the better approach.
What Kubernetes-native workflows should this support?
The primary value is providing a Kubernetes-native control plane for OpenShell. With CRDs for sandboxes, policies, and providers, native K8s RBAC applies directly:
- Tenant Admin controls which policies and providers are available, who can create sandboxes within their domain
- Platform Admin scopes Tenant Admins via namespace-level RBAC
- No separate authorization layer required
Workflows that follow naturally:
- Declarative sandbox lifecycle from GitOps tools, CI pipelines, or any K8s API client — no orchestration layer needed beyond
kubectl apply - Policy-as-code (providers and policies are a natural fit for declarative spec — they change infrequently and benefit from version control)
The operator should also make minimal assumptions about the underlying infrastructure — particularly around networking and security. Enterprise K8s clusters vary widely in how traffic enters and exits (Ingress, Gateway API, NodePort, LoadBalancer), whether a service mesh is present (Istio, Linkerd, mTLS between services), how network policies are enforced, and what cluster-level security standards apply (PodSecurityStandards, AppArmor/seccomp defaults). Sandboxes — or at least the proxy component — need to talk to other cluster services and external APIs, and must play well with existing network and security infrastructure.
Today this is complicated by the secondary network namespace architecture — service mesh sidecars expect a single network namespace per pod, which conflicts with the current model. In that context, #1305 (removing the secondary network namespace and splitting the proxy from the supervisor/agent) is important — it's what makes coexistence with meshes and standard NetworkPolicy infrastructure possible.
In a Kubernetes deployment, the Sandbox Consumer interacts through
kubectland the K8s API. The TUI and direct gRPC should not be required.What should the first version include?
At minimum:
- Installation and upgrades of the gateway and the OpenShell control plane
- NetworkPolicy generation for gateway and sandbox pods
- Agent-sandbox-controller installation (could be externalized — it has its own release lifecycle — but debatable)
However — deploying the gateway in K8s without additional native capabilities is limiting. A Kubernetes-native frontend for sandbox lifecycle management is necessary for the operator to deliver value beyond what Helm already provides.
This connects to the boundary question posed in the issue. In my view the operator is:
A replacement for some gateway responsibilities inside Kubernetes environments.
Specifically, it replaces the gateway's control-plane half. The gateway today conflates:
- Control-plane: providers, policies, sandbox lifecycle CRUD, persistence
- Data-plane: supervisor sessions, relay byte-streaming, exec/forward/SSH, service routing
These are fundamentally different concerns — and Kubernetes already has a native separation between them. The operator absorbs the control plane:
- CRDs become the persistence layer (etcd replaces SQLite/Postgres)
- RBAC replaces custom authorization
- Reconciliation loops replace the gateway's CRUD endpoints
The gateway shrinks to its data-plane role — relay, streaming, supervisor sessions — and becomes a slimmer runtime service that the operator deploys and manages.
This mirrors how Kubernetes itself separates concerns: API server + controllers own declarative state; kubelets and kube-proxy own runtime execution. An OpenShell operator should follow the same pattern.
The Sandbox CRD shape depends on how the
agents.x-k8s.iorelationship resolves (see below), so it could be phased slightly later. Policies and Providers are natural first-class CRDs from the start — the Tenant Admin persona needs them immediately.Should proposed sandbox resources be user-facing, platform-team-facing, or both?
Both, layered by persona:
- Sandbox resource — Sandbox-Consumer-facing. RBAC controls who can create them.
- Policy resource — Tenant-Admin-facing. Defines the security boundary sandboxes must reference.
- Provider resource — Tenant-Admin-facing. Makes credentials available within a trust domain.
- Gateway resource — Platform-Admin-facing. Defines the deployment.
The Sandbox Consumer picks from pre-approved policies and providers by name. The Tenant Admin authors those. The Platform Admin doesn't touch either.
This is the platform-centric multi-tenant view. In single-player or developer-centric deployments (one person wearing all hats), the rigid separation is optional — but the system should still support:
- Tenant Admin policies as enforceable baselines (even in single-player, a default policy should apply)
- Self-service within those bounds (the developer can configure their own sandbox without a separate admin approving each request)
Key point: the gateway should not be the sole entrypoint for sandbox creation. In enterprise deployments where frontend services are separate systems, those services should create sandboxes directly via the K8s API — as simple as creating a CR.
How should a proposed OpenShell sandbox resource relate to the existing Agent Sandbox CRD?
This is the most difficult design question and relates directly to #1680.
The tension:
- ASB resources have capabilities (warm pools, scale-to-zero) that OpenShell will increasingly want to leverage
- OpenShell exposes concepts (policies, inference routing, provider attachment, audit) with no representation in the ASB CRD
- Every new ASB feature will require integration work; every OpenShell-specific field has no upstream home
The exact layering needs careful design. What is clear: the user-facing API should be OpenShell's, not the raw ASB primitive.
What fields would you expect on an OpenShell sandbox custom resource?
Spec:
gatewayRef— which gateway manages this sandboxpolicyRef— reference to a reusable policy resource (see policy question below)image— sandbox container imageentrypoint— override command (immutable after creation)environment— environment variables for the supervisorproviders— list of Provider names to attachgpu— GPU resource request
Pod-level fields (runtimeClassName, imagePullSecrets, resource requests, published ports, volumes) depend on the
driver_configdesign in #1589 landing first — those should follow whatever shape that takes.Status:
phase— Pending / Provisioning / Ready / Error / DeletingsandboxID— UUID assigned by the gatewayconditions— e.g. Synced, GatewayReachable, PolicyLoaded, UpstreamSandboxReadyobservedGeneration— required for GitOps health checks (ArgoCD/Flux use this)
How should policies be referenced or embedded?
Referenced, not embedded. A dedicated reusable policy resource that sandboxes point to by name. The Tenant Admin defines the security boundary (filesystem access, process identity, network rules); the Sandbox Consumer just references it.
Benefits:
- Tenant Admins manage policies independently of Sandbox Consumers
- Policy updates propagate to running sandboxes for mutable fields (network rules)
- Validation webhook prevents deletion of a policy still referenced by active sandboxes
- Per-field immutability constraints: process identity is immutable after creation; network rules are freely updatable; filesystem paths are additive-only
- Opens the door for the policy advisor to surface suggestions as proposed modifications
How should credentials, provider configuration, and inference routing be handled?
A dedicated Provider CRD referencing a Kubernetes Secret:
spec.type— provider type (anthropic, openai, nvidia, etc.), immutablespec.credentialsSecretRef— references a Secret (all data keys become credential entries)spec.config— non-secret configuration (model overrides, base URLs)
This is a Tenant Admin resource. The Sandbox Consumer references providers by name — never sees credential values, never touches Secrets directly. The operator watches the referenced Secret and re-syncs on credential rotation without requiring a Provider CR edit. Standard K8s secret management (external-secrets-operator, vault-agent) integrates without additional work.
Are there existing operator patterns or tools we should align with?
controller-runtime with standard patterns:
- Finalizer-based lifecycle for guaranteed cleanup
CreateOrUpdatefor idempotent reconciliation- Config hash annotations on pod templates for rolling restarts on config changes
- Leader election for the cluster-wide operator
- Watches on owned resources for immediate reaction (not polling)
- Compatible with existing cluster infrastructure: security standards (PodSecurityStandards, AppArmor, seccomp), common upstream controllers (cert-manager), and networking/service-mesh integrations
Additional Recommendations
-
Naming: Do not reuse the kind
Sandboxin the same cluster asagents.x-k8s.io/Sandbox. UseOpenShellSandboxor a distinct API group. Tooling confusion (kubectl, k9s, ArgoCD resource trees) is significant in practice. -
Hardened nodes: The K8s compute driver needs to set AppArmor and seccomp profiles on sandbox pods, not just capabilities. On hardened distributions (Garden Linux, Bottlerocket), sandbox creation fails without
Unconfinedprofiles due to network namespace setup. Related: #1650 is important — removingCAP_SYS_ADMINis the correct direction for secure K8s deployments. -
Entrypoint: Expose command/entrypoint in the sandbox creation API. Hardcoding
sleep infinityforces admission webhook workarounds for CI and batch workloads. -
driver_config is mandatory: Private registries (imagePullSecrets) and alternative runtimes (runtimeClassName for GPU/kata/gVisor) are standard in enterprise K8s. These need first-class representation in the compute driver — #1589 is the right direction and a prerequisite for a serious operator.
-
Pod selection contract: Document the
agents.x-k8s.io/sandbox-name-hashlabel as stable API. Operators writing NetworkPolicies for sandbox pods need a reliable, documented selector — replicating internal hashing logic is fragile.
A Note on Deployment Profiles
The control-plane / data-plane separation above naturally leads to two deployment profiles:
-
Gateway-fronted (resembling the current k3s setup): the gateway exposes the user-facing API (TUI/CLI/SDK over gRPC). The operator handles deployment and upgrades but does not replace the gateway's control-plane role. The gateway retains both control-plane and data-plane responsibilities.
-
Operator-fronted (OpenShell as an embeddable system): the operator/CRDs are the user-facing API. Frontends and services create Sandbox CRs via the K8s API. The gateway is reduced to data-plane only (relay, streaming, supervisor sessions). The Sandbox Consumer (often an automated system, not a human) creates Sandbox CRs directly.
The operator design could try to accommodate both. The difference is how far the control-plane/data-plane decomposition goes — in gateway-fronted mode the gateway keeps both hats; in operator-fronted mode the operator takes the control-plane hat entirely.
Reacted by Ron LevFeedback from a user here planning on utilising a sizeable amount of openshell + nemoclaw sandboxes in our company (Platform Engineer)
I would really appreciate having an openshellsandbox CRD natively in the controller, in my org we have adopted nemoClaw + Openshell and we had to make a wrapper controller around it to achieve nice gitops workflow to spawn up instances.
continuing to utilize agent-sandbox under the hood is fine, as its an official SIGS project thus will align with industry standards
I made a draft MR above implementing this feature and it works when testing but I guess this topic would involve a lot more discussion and probably several tweaks/improvements to this MR.
I am happy to get involved in this and potentially be a contributor.
I have made a vouch request
#1829Reacted by Sanket Nadkarni@drew a couple of answers/recommendation for your questions. The use cases we are looking for is enterprise production deployment of OpenShell on Kubernetes and integration with NVIDIA Blueprints at scale. At high level and for first implementation, OpenShell should use a Kubernetes-native controller-gateway design. The OpenShell Kubernetes deployment should run one OpenShell control-plane service that has two roles:
- It acts as the Kubernetes controller/operator.
- It acts as the OpenShell gateway/runtime API.
Of course there are pros and cons for this design mainly API endpoint scale at a different rate with different trigger metrics, but it simplifies the architecture as opposed to having one micro service for the operator and one for the gateway and many for the agent sandboxes. There are many example operators in the industry that adopted this pattern: Argo CD, Prometheus Operator and cert-manager.
Should the operator deploy OpenShell, manage sandboxes, or both?
Both, in one OpenShell controller-gateway service.
The Kubernetes deployment should run a single OpenShell control-plane service that:- Acts as the Kubernetes controller/operator
- Acts as the OpenShell gateway/runtime API
It should watch OpenShell sandbox CRDs such as:
apiVersion: openshell.ai/v1alpha1 kind: OpenShellSandboxWhat Kubernetes-native workflows should this support?
It should support a simple YAML-first workflow centered on one main CRD:kind: OpenShellSandboxA user should be able to create a sandbox with one Kubernetes object:
apiVersion: openshell.ai/v1alpha1 kind: OpenShellSandbox metadata: ......The OpenShell controller-gateway should:
watch OpenShellSandbox objects validate requested image/resources/providers/policy deny requests the user or namespace is not allowed to use create agents.x-k8s.io/Sandbox serve runtime gateway APIs for the sandbox update OpenShellSandbox status clean up child resources on deletionMulti-tenancy should use normal Kubernetes controls:
namespaces RBAC ResourceQuota LimitRange NetworkPolicy admission validationWhat should the first version include?
The first version should be small and focused with one micro service installed by helm that combines both gateway endpoint function and operator.Should OpenShell sandbox resources be user-facing, platform-team-facing, or both?
Both. OpenShellSandbox should be usable by both platform teams and application/user teams, but with different permissions and policy boundaries.
Platform teams should be able to create sandboxes for managed services, shared demos, CI systems, etc
Application/user teams should be able to create sandboxes in namespaces where they have permission.
The difference is what each persona is allowed to do.Platform team: can install/configure the controller-gateway can grant RBAC can set namespace quotas and policies can define allowed registries, providers, resources, and runtime settings can create OpenShellSandbox objects for shared/platform workloads Application/user team: can create OpenShellSandbox objects only in allowed namespaces can use only allowed images, resources, providers, policies, and secrets cannot change controller-gateway configuration cannot bypass namespace or admission policyIf a user requests something outside their allowed boundary, the controller-gateway should reject it or mark it rejected:
status: phase: Rejected conditions: - type: Accepted status: "False" reason: ProviderNotAllowed message: "Provider ai-inference is not allowed in namespace team-a"How should OpenShellSandbox relate to the Agent Sandbox CRD?
OpenShellSandboxshould be the OpenShell-facing CRD abstractingSandboxCRD. Each operator reconciles its own CRD.agents.x-k8s.io/Sandboxshould not be exposed by OpenShell operator.
Tradeoff: OpenShell must stay compatible with Agent Sandbox. That means OpenShell must track Agent Sandbox CRD versions and schema changes and track the api version. Which also means deployment and runtime api version checking and compatibility.How should policies be referenced or embedded?
While embedded is easier to understand and require fewer CRDs, it makes more sense to use reference when resource parameters grow large.How should credentials, provider configuration, and inference routing be handled?
For v1, keep them inline in OpenShellSandbox, but never embed raw secret values.
Use Kubernetes Secret references for credential material.
Example:apiVersion: openshell.ai/v1alpha1 kind: OpenShellSandbox metadata: name: teama-openclaw namespace: team-a spec: image: <image-url> providers: - name: nvidia-inference type: openai credentialSecretRef: name: teama-api-key key: api-key config: baseUrl: https://integrate.api.teama.com/v1The controller-gateway should read only referenced Secrets in the sandbox namespace and validate the user is allowed to reference those Secrets.
We're looking at Kubernetes-native integration from the perspective of platform teams deploying OpenShell as part of larger agent orchestration systems. Our interest is in what the
OpenShellSandboxCRD could represent beyond infrastructure lifecycle. This is the real tension: What actually is the scope of OpenShell itself ?The Boundary Is Already Blurry
The existing comments frame the operator as managing sandbox infrastructure: images, resources, lifecycle, policies. But OpenShell itself already crosses the line between agent-agnostic infrastructure and agent-aware configuration in several places. Provider credentials aren't simple env var passthrough; they use placeholder-based injection with multiple refresh strategies, proxy-level credential rewriting, and provider profiles categorized by agent concern (
INFERENCE,AGENT,SOURCE_CONTROL,MESSAGING). Inference routing viainference.localtransparently intercepts agent API calls and routes through centrally managed model configuration. The proxy understands MCP at a semantic level and enforces per-tool-name policies. A built-in skill system installs policy advisor capabilities into sandboxes, teaching agents to self-manage network policy through structured proposals with formal verification.OpenShell already knows it's running agents, not just containers. The question for the operator is whether the Kubernetes CRD should reflect that, or stop at the infrastructure boundary and leave agent-aware configuration to the gateway API. If OpenShell doesn't provide a declarative Kubernetes surface for this, out-of-tree platforms will build their own CRDs on top of OpenShell's primitives. They will do that regardless. The question is whether OpenShell wants to shape that space.
What the Extended CRD Could Look Like
If the scope extends into agent configuration, the CRD becomes the declarative surface for the full OpenShell experience:
Skills: Declared on the CR, referencing a ConfigMap, a skill catalog, or inline definitions. The operator gets skills into the sandbox (mount, upload via gateway API, keep in sync on updates). This is the core declarative-to-imperative bridge problem.
MCP server connections: Remote MCP servers (SSE, Streamable HTTP) declared with endpoint URLs and credential references. These are independent of the sandbox image since the connection goes over the network through the proxy. The operator translates these into gateway-side MCP policy rules and handles credential rotation by watching the referenced Secrets.
Observability wiring: OTel collector endpoints, tracing configuration, and environment variables declared on the CR rather than scattered across Helm values and ConfigMaps.
Status as discoverability: If OpenShell integrates with A2A eventually, the agent card (metadata, capabilities, endpoint URL) could be exposed in
.status, making sandboxes discoverable by other agents through standard Kubernetes API queries.apiVersion: openshell.ai/v1alpha1 kind: OpenShellSandbox metadata: name: my-agent namespace: team-ml spec: gatewayRef: name: default-gateway policyRef: name: team-default-policy image: nvcr.io/nvidia/openshell:latest # Agent capabilities (if scope extends beyond infrastructure) skills: - name: web-search sourceRef: configMapRef: { name: skill-web-search } - name: code-review catalogRef: { name: corporate-catalog, version: "1.2" } # Remote MCP servers (no sandbox image dependency) mcpServers: - name: github transport: streamable-http url: https://mcp.github.internal/sse credentialRef: secretRef: { name: github-token, key: token } - name: jira transport: sse url: https://mcp-gateway.corp.internal/jira credentialRef: secretRef: { name: jira-oauth, key: access-token } observability: tracing: endpoint: otel-collector.monitoring:4317 environment: - name: PROJECT_ID value: alpha - name: API_KEY valueFrom: secretKeyRef: { name: my-secret, key: value } status: phase: Ready sandboxID: abc-123 conditions: - type: SkillsSynced status: "True" - type: MCPConnected status: "True" agentCard: name: my-agent capabilities: ["code-execution", "web-search"] url: "https://..." observedGeneration: 3
Multi-Tenancy and the Multi-Player RFC
This design should be considered alongside the multi-player RFC (#1980). Both Kubernetes namespace modes from that RFC (managed and operator) map OpenShell workspaces to K8s namespaces, differing only in who provisions the namespace.
For control plane operations (who can create/delete
OpenShellSandboxCRs, manage providers and policies), K8s RBAC is a natural fit. For data plane operations (exec into a running sandbox, stream relay output, share sessions), the gateway needs its own authorization model since K8s RBAC doesn't reach into gRPC streaming endpoints. The multi-player RFC's role model covers this. Both are needed, with clear boundaries.The Declarative-to-Imperative Bridge
If the operator extends into agent configuration, the hardest engineering problem isn't deployment or lifecycle. It's continuously reconciling declarative CRD state with the gateway's imperative API. Skills change, MCP connections come and go, credentials rotate, policies update. The operator must detect drift between the CR's desired state and what the gateway actually has configured, and reconcile without disrupting running sessions.
This is what would make the operator more than a Helm replacement: an active reconciliation controller translating Kubernetes-native desired state into gateway API calls, continuously.
Answers to the Specific Questions
Would you use an operator mainly to deploy OpenShell, to manage sandboxes, or both?
Both. Deployment is table stakes. The value question is how far sandbox management extends: infrastructure only, or agent configuration too.What Kubernetes-native workflows should this support?
At minimum: declarative sandbox lifecycle via GitOps. If scope extends: skills, MCP connections, credentials, observability, and policies all declared on the CR and reconciled by the operator.What should the first version include?
Gateway deployment,OpenShellSandboxCRD with policy and provider references, basic status reporting, and the reconciliation loop against the gateway API. Agent configuration fields (skills, MCP) could follow once the bridge pattern is proven.Should proposed sandbox resources be user-facing, platform-team-facing, or both?
Both, scoped by K8s namespace RBAC for control plane operations.How should a proposed OpenShell sandbox resource relate to the existing Agent Sandbox CRD?
OpenShellSandboxis a higher-level abstraction that generates Agent Sandbox resources under the hood. Users never interact withagents.x-k8s.io/Sandboxdirectly. +1 on usingOpenShellSandboxto avoid tooling confusion.What fields would you expect?
See the spec sketch above. Infrastructure fields (image, resources, policyRef, gatewayRef, environment) are clear. Agent configuration fields (skills, mcpServers, observability, agentCard in status) depend on scope decisions.How should policies be referenced or embedded?
Referenced.How should credentials, provider configuration, and inference routing be handled?
Providers as configuration resources referencing K8s Secrets. The operator resolves Secret references at runtime and watches for rotation. Never stores raw credential material on the CR.What status, events, and observability would be required?
Phase, conditions,observedGeneration,sandboxID. If agent configuration is in scope:SkillsSynced,MCPConnectedconditions, and optionallyagentCardfor A2A discoverability.Are there existing operator patterns or tools we should align with?
controller-runtime with standard patterns. If agent configuration is included, the reconciliation loop needs reliable drift detection against the gateway's imperative API, which is more complex than typical CRD-to-resource controllers.- added a commit that references this issue
on Jul 17, 2026 We've built this as an out-of-tree operator, now usable: https://github.com/lensapp/openshell-k8s-operator. Feedback and testers welcome.
The design choice that answers most of the disambiguation questions here: it's a thin front-end, not a second control plane. CRDs in a distinct group (openshell.lenshq.io/v1alpha1) — OpenShellSandbox, OpenShellProvider, OpenShellPolicy, OpenShellWorkspace — translate into gateway gRPC calls and mirror gateway state into .status. So:
- Separate CRDs, not agents.x-k8s.io/Sandbox re-exposed; we never reuse the kind name Sandbox.
- The gateway stays the source of truth — its compute driver still owns the agents.x-k8s.io/Sandbox object.
- Additive, not either/or: the operator authenticates to the gateway as an OIDC client; the Helm chart bundles gateway + issuer + operator, or targets a BYO gateway.
It maps onto the three-persona model @kon-angelo described, via native RBAC. In testing it also grew a webhook that confines kubectl exec into the privileged sandbox pods, plus CEL immutability on identity fields. It reuses OpenShell's own crates pinned to an exact rev (v0.0.90), so it tracks upstream rather than forking.
On the in-tree / controller-gateway direction (@nakfour): no stake in staying out-of-tree. Keeping the gateway as authority made the operator thin; if OpenShell converges on an in-tree operator, happy for this to be a reference, a migration source, or superseded.
Reacted by zer0tweetsPerspective from the OpenShift / OLM packaging side. Most of the thread has debated what the operator does; I want to add the constraint set that shows up when you try to ship it as an OperatorHub-installable operator, because it turns out to disambiguate a couple of the open questions rather than just adding requirements on top of them.
To be clear up front: I am not proposing a separate "OpenShift operator". OpenShift is a packaging and admission target for whatever operator this issue converges on. But the packaging constraints are load-bearing enough that I think they belong in the design discussion rather than after it.
1. OLM has an opinion on the control-plane / data-plane question
@nakfour proposes one service that is both controller and gateway/runtime API, citing Argo CD, Prometheus Operator, and cert-manager. Worth noting that all three of those are counter-examples to the merged shape: the Prometheus Operator does not become Prometheus, cert-manager's controller is separate from its webhook and CA injector, and Argo CD's operator deploys the Argo components rather than serving the Argo API itself.
That split is not stylistic under OLM. The operator's Deployment is created and owned by OLM from the ClusterServiceVersion
install.spec.deployments, which means:- OLM controls its replica count, and upgrades roll it by replacing the CSV. A byte-streaming data plane (relay, exec, forward, SSH) being torn down by a catalog-driven operator upgrade is a materially different failure mode than a reconciler being restarted.
- The operator scales on the number of reconciled objects. The gateway data plane scales on concurrent sandbox sessions and bytes. Fusing them puts two unrelated scaling curves in an OLM-managed Deployment that the cluster admin is not supposed to hand-edit.
- Leader election is expected for the controller and wrong for the data plane.
This doesn't decide how far the control-plane decomposition goes — @kon-angelo's gateway-fronted vs operator-fronted framing still stands. It does mean that whichever direction wins, the reconciler and the streaming data plane should be separate Deployments. @jakolehm's thin front-end shape is the one that survives this constraint without further work, and the operator-fronted variant survives it too as long as the shrunken gateway stays its own workload.
2. Agent Sandbox as an OLM dependency
sandboxes.agents.x-k8s.iois installed today by applying the upstream manifest (docs/kubernetes/setup.mdx:36). OLM cannot express "depend on a CRD owned by a controller that isn't packaged as an operator" —requiredCRDs in a CSV resolve against other bundles in a catalog, not againstkubectl apply. So there are exactly three options and it's worth picking one deliberately:- Declare
sandboxes.agents.x-k8s.ioas arequiredCRD and document that Agent Sandbox must be installed first. Install fails with a resolution error until it is. Honest, but a poor OperatorHub experience. - Ship an OpenShell-maintained bundle of the Agent Sandbox controller and declare a real OLM dependency. Means taking on its release cadence, which @kon-angelo already flagged as debatable.
- Have the operator install and manage the Agent Sandbox controller itself. Conflicts with clusters that already run it for other workloads.
I'd argue for (1) in v1 with a clear precondition check surfaced as a status condition, but the choice needs to be explicit because it's visible to every OperatorHub user on day one.
3. The current OpenShift path cannot be packaged as-is
This is the concrete blocker, and it's independent of the CRD debate.
docs/kubernetes/openshift.mdxdocuments the supported OpenShift install as: pre-create the namespace,oc adm policy add-scc-to-user privileged -z openshell-sandbox, and--set server.disableTls=true. It carries an explicit experimental warning.An operator whose sandboxes require the
privilegedSCC will not pass Red Hat operator certification, and in most enterprise OpenShift clusters it will not get past the platform team regardless of certification.privilegedis the SCC of last resort; requesting it for the workload pods rather than the operator is a hard stop in a lot of shops.Looking at what the driver actually asks for,
privilegedlooks broader than necessary:- Combined topology requests
SYS_ADMIN, NET_ADMIN, SYS_PTRACE, SYSLOGon the agent container (crates/openshell-driver-kubernetes/src/driver.rs:3572). - Sidecar topology (
SupervisorTopology::Sidecar,crates/openshell-driver-kubernetes/src/config.rs:66) puts network supervision in a sidecar requestingNET_ADMIN, NET_RAW, CHOWN, FOWNER(driver.rs:2846) and leaves the agent container as a low-capability wrapper. - The AppArmor profile is already configurable (
RuntimeDefault/Unconfined/Localhost/<name>), which is the right shape for hardened nodes — this addresses part of @kon-angelo's point 2.
That reads like a named custom SCC shipped by the operator, not
privileged: a small capability set,RunAsAnyfor UID/fsGroup so OpenShift admission still assigns them, and the sidecar topology as the OpenShift default so the blast radius sits in the proxy sidecar rather than the agent container. Two things would need resolving to get there: the combined-topologyCAP_SETPCAPgap in #2751, and the hardcoded UID/fsGroup that #2697 and #2707 already cover.The
server.disableTls=truerequirement is the other half. PassthroughRoutesupport already exists in the chart (openshiftRoute.*,values.yaml:523) and #2466 landed cert-manager external issuer support, so the TLS-on path is reachable today — it just isn't the documented default. An operator should only ever install the TLS-on configuration.I think this is the real v1 prerequisite, ahead of any CRD work. If OpenShell runs on OpenShift only under
privilegedwith TLS off, the operator has nothing certifiable to package no matter which CRD shape wins.4. Disconnected installs constrain the CRD, not just the bundle
docs/kubernetes/setup.mdx:42currently describes air-gapped installs as a manual mirroring exercise across four images: the Agent Sandbox controller,image.repository(gateway),server.supervisorImage, andserver.sandboxImage.OLM's answer is
relatedImagesin the CSV with every image pinned by digest, sooc-mirrorcan resolve the full closure andImageDigestMirrorSetcan rewrite them. The gateway, supervisor, and controller images are all pinnable that way. The sandbox image is not — it's per-sandbox and user-supplied, so it can never appear inrelatedImages.That has a CRD consequence worth deciding now rather than later: a disconnected cluster needs the operator to validate sandbox images against an admin-controlled allowlist and to surface a clear rejection when a user references an unmirrorable image. That's a Tenant-Admin-facing concern in @kon-angelo's model and it argues for image policy living on a platform-owned resource rather than being free-form on the Sandbox CR. It also lines up with @nakfour's
Rejectedphase /Accepted: Falsecondition suggestion —ImageNotMirroredis exactly that shape.5. Install modes map onto the three-persona model
@kon-angelo's Platform Admin / Tenant Admin / Sandbox Consumer split translates cleanly onto OLM install modes, which is a useful consistency check on the CRD layering:
AllNamespaces— Platform Admin installs the operator once;Gateway(or equivalent) is cluster-scoped or lives in the operator's namespace.OwnNamespace/SingleNamespace— Tenant Admin ownsPolicyandProviderin their namespace.- Sandbox Consumer creates namespaced Sandbox CRs, gated by native RBAC.
If any of the proposed resources can't be cleanly assigned to one of those tiers, that's a signal the layering isn't right yet. Concretely, it means Policy and Provider should be namespaced with an optional cluster-scoped default, not cluster-scoped only.
6. Capability levels as shared vocabulary for "what should v1 include"
The question "what should the first version include" has had several good answers with no common yardstick. The Operator Framework capability levels are one, and they're what OperatorHub displays:
- L1 Basic Install — gateway + dependencies installed and configured. This is roughly Helm parity and, per the thread's consensus, not enough on its own.
- L2 Seamless Upgrades — version upgrades handled, which is where OLM earns its place over
helm upgrade. - L3 Full Lifecycle — sandbox/policy/provider lifecycle, backup of gateway state.
- L4 Deep Insights — metrics, alerts, status conditions. Connects to feat(observability): OpenTelemetry export surface for gateway operators #2507 and feat(server): metrics instrumentation #909.
- L5 Auto Pilot — auto-scaling, auto-remediation, auto-tuning.
L2 is the honest floor for shipping to OperatorHub at all, and L3 for the declarative sandbox workflow this thread is actually asking for. Naming a target level would make "what's in v1" concrete without pre-committing to a CRD shape.
Suggested sequencing
- Make OpenShell installable on OpenShift under a named custom SCC with TLS enabled — resolving Kubernetes driver never requests CAP_SETPCAP, so combined-topology sandboxes crash-loop with EPERM on privilege drop #2751, feat(helm): expose sandbox_uid / sandbox_gid as first-class Helm values #2697, Support low numeric UID and GID values. #2707 and promoting the sidecar topology as the OpenShift default. Independent of the operator debate, and a prerequisite for any certifiable bundle.
- Decide the Agent Sandbox dependency posture (§2), since it's user-visible from the first install.
- Whatever operator shape wins, keep the reconciler and the streaming data plane as separate Deployments (§1).
- Add image-policy validation to the platform-facing resource so disconnected clusters work (§4).
No strong preference between in-tree and out-of-tree — the SCC and TLS work in (1) is needed either way.
@flg77 thanks for your input, but you add a lot of OpenShift specific concerns that goes beyond a vanilla Kubernetes operator (the topic of this issue). I think its worth to create a separate issue for OpenShift specific concerns, like the OpenShift installation issues you address (which are only slightly connected to the operator discussion at hand). Your comment is also a bit hard to consume by humans, maybe you can summarize in your own words, what an pure Kubernetes Operator for the control and data plane should look like ? That would be very helpful. We are still in an early discussion phase, so we are not yet there for a concrete implementation plan. Let's keep the discussion on a personal level (agent assisted, but with a human focus, and with some empathy for the readers).
Reacted by Francisco Javier ArceoHi @drew - I realize the issue title is "Kubernetes Operator" but it seems like an immediate gap for K8s installations is a CRD for an OpenShell Sandbox with a controller for lifecycle management. As one of the comments suggests, the controller can be a thin layer above the existing API. Would it make sense to split that out and tackle it separately from an operator that installs and manages OpenShell? Is an RFC the right next step?
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsIdea
Problem Statement
OpenShell's Kubernetes deployment path currently uses the gateway as the user-facing control plane and depends on the Kubernetes Agent Sandbox controller for runtime sandbox pods. As Kubernetes users adopt OpenShell, we need to decide whether an OpenShell Kubernetes Operator should exist, what responsibilities it should own, and how it should relate to the existing gateway and Agent Sandbox CRD.
This matters because platform teams may expect Kubernetes-native installation, reconciliation, status, and declarative sandbox workflows. At the same time, the gateway already owns OpenShell API behavior, credentials, inference configuration, policy attachment, sandbox lifecycle APIs, logs, watch streams, and client integrations through the CLI, SDKs, and TUI. An operator could complement that model, overlap with it, or replace parts of it in Kubernetes environments.
We should collect user, platform team, and contributor feedback before committing to a specific operator shape.
Proposed Design
Explore an OpenShell Kubernetes Operator with two possible directions:
OpenShell Deployment Operator
The operator could install and manage an OpenShell gateway inside a Kubernetes cluster. It may handle:
OpenShell Sandbox Custom Resource Operator
The operator could introduce an OpenShell-owned custom resource for declarative sandbox management. This resource may let users or platform teams define:
.status.Existing Agent Sandbox CRD Boundary
OpenShell already depends on the Kubernetes Agent Sandbox CRD in Kubernetes deployments today. That resource is also named
Sandbox, but it lives in theagents.x-k8s.ioAPI group and is reconciled by the Agent Sandbox controller into Kubernetes pods.The existing Agent Sandbox CRD is a runtime-level Kubernetes primitive used by the OpenShell Kubernetes compute driver. A proposed OpenShell sandbox custom resource may be a different API surface: a user- or platform-facing OpenShell object that includes OpenShell policy references, inference configuration, provider wiring, lifecycle semantics, audit expectations, and gateway integration.
The design should disambiguate these two concepts before proposing a CRD shape:
agents.x-k8s.io/Sandboxwith any OpenShell-specific sandbox resource?Operator And Gateway Boundary
The operator and current OpenShell gateway likely have overlapping responsibilities, especially around sandbox lifecycle, configuration, status, policy attachment, and infrastructure integration. It may be valid for some deployments to use an operator or a gateway, but not both.
The design should clarify when users interact with the Kubernetes Operator directly versus when they use the existing gateway API, CLI, SDKs, or TUI. In particular, we should decide whether the operator is primarily:
Questions For Feedback
Non-Goals For Now
Alternatives Considered
agents.x-k8s.io/SandboxCRD directly as the user-facing sandbox API. This avoids a second sandbox resource, but may leak a runtime-level primitive that does not model OpenShell policy, credentials, inference routing, audit behavior, or gateway integrations.Agent Investigation
The local repo already documents and uses the existing Agent Sandbox CRD:
deploy/kube/manifests/agent-sandbox.yamldefinessandboxes.agents.x-k8s.iowithkind: Sandbox.deploy/helm/openshell/README.mdsays the Kubernetes Agent Sandbox CRDs and controller must be installed before deploying OpenShell.architecture/gateway.mddescribes the gateway as the OpenShell control plane and notes that Kubernetes sandbox authentication verifies the pod's controllingSandboxownerReference against the live Sandbox CR UID.crates/openshell-driver-kubernetes/contains the Kubernetes compute driver that creates and watches Agent Sandbox resources.These findings make the operator design a boundary and API-shape question, not just a request to add a new controller.
Proposed Outcome
Use this issue to gather requirements and decide whether to create one or more follow-up design issues for: