Kubernetes as the Agent Control Plane in 2026: Agent Sandbox Hit v1.0, 300 Claims a Second, and DRA in Every Supported Release
I run agents in Firecracker microVMs outside Kubernetes, on PandaStack, the Apache-2.0 microVM cloud I build and operate, so weigh my affiliation when I talk about the alternatives. The question platform teams ask me most this year is not "which sandbox vendor" but "can we do this on the cluster we already have". A year ago the honest answer was: sort of, with a StatefulSet of size one, a headless Service, a PersistentVolumeClaim and a RuntimeClass you wired up yourself. In October 2026 the answer is different, because Kubernetes grew an agent-shaped API, the device scheduling model it needed went stable in every supported release, and the framework and gateway layers above it stopped being conference demos. This is the ledger of what shipped, with versions and dates, where the numbers come from, and where I think the cluster still loses to a dedicated microVM fleet.
The substrate: four releases and one GA that matters
Kubernetes 1.34 landed on 27 August 2025 with 58 enhancements, 1.35 on 17 December with 60, 1.36 on 22 April 2026 with 70, and 1.37 on 26 August with 67, of which 16 went stable, 23 beta and 27 alpha. The releases page has 1.34 going end-of-life on 27 October 2026, which means that from this month every supported Kubernetes ships Dynamic Resource Allocation as a GA API.
That is the one graduation an agent platform cares about. DRA's resource.k8s.io/v1 ResourceClaim, ResourceClaimTemplate and DeviceClass went GA in 1.34, and the 1.35 notes say the feature gate is now always enabled. The 1.36 update made the prioritized list stable, so a claim can say "an H100, or failing that an A100", and moved partitionable devices, device taints, binding conditions and resource health to beta. The 1.37 update graduated extended-resource support, so a Pod that still asks for example.com/gpu is satisfied by a DRA driver with no device plugin alongside it, made device taints stable, standardised a numaNode attribute, and added alpha CEL-derived attributes and compatibility groups that stop the scheduler pairing a MIG slice with a vGPU profile on the same card. NVIDIA's DRA driver reached v0.5.0 on 19 August. For agents the consequence: a sandbox that needs an accelerator asks for a device by attributes, and a partitionable device means a MIG slice per sandbox rather than a whole card.
The demand side is in the CNCF's 2025 annual survey, published 20 January 2026: 82 percent of container users run Kubernetes in production, up from 66 percent in 2023; 66 percent of organisations hosting generative AI models use Kubernetes for some or all of their inference; 44 percent do not yet run AI or ML workloads on it at all. The CNCF's Kubernetes AI Conformance repository held 69 submissions across versions 1.33 to 1.37 when I counted on 5 October, from EKS, GKE, AKS and OpenShift down to Talos and k0s. KubeCon Europe ran in Amsterdam on 23 to 26 March; KubeCon North America is in Salt Lake City on 9 to 12 November with a new AI Inference and Agentic track and a Tim Hockin session titled "Kubernetes Solutions for Agent-Shaped Problems". It has not happened yet.
Agent Sandbox: a preview in November, v1.0 in August
The project that turned "sort of" into "yes" is kubernetes-sigs/agent-sandbox, a SIG Apps subproject under Apache-2.0 with 4,142 stars. The repository was created on 12 August 2025 and Google announced it on 11 November at KubeCon Atlanta, with gVisor and Kata as the isolation runtimes and a warm-pool orchestrator for sub-second starts. The Kubernetes blog's March 2026 post by Janet Kuo and Justin Santa Barbara defines the workload the project targets more carefully than most vendors do: isolated, stateful, singleton, mostly idle, needing a stable identity and a suspend-and-resume lifecycle. It also puts a plain Pod's cold start at about a second, fine for a rollout, not for an agent woken mid-conversation.
The release history is the stability signal. v0.5.0 on 24 June 2026 moved the API to v1beta1; v1.0.0 on 28 August removed v1alpha1, added the in-pod sandboxd daemon and browser routing; since then it has shipped weekly, with v1.0.2 adding Ed25519 execution-scoped tokens, v1.0.4 adding lifecycle events, and v1.0.5 on 1 October adding native in-cluster sandboxd transport and TypeScript runtime connectivity. There are four custom resources. Sandbox in agents.x-k8s.io/v1beta1 wraps a pod template with an operatingMode of Running or Suspended, an optional shutdownTime, and a stable serviceFQDN in its status. The extensions group extensions.agents.x-k8s.io/v1beta1 adds SandboxTemplate, SandboxWarmPool (replicas plus a template reference) and SandboxClaim, which hands a caller a pre-warmed pod from the pool. Isolation is delegated to a RuntimeClass, and the repository's Firecracker example runs the same template on Kata's kata-fc runtime, claiming roughly 125 ms boots, with the honest caveats that nodes need KVM and containerd needs a devmapper snapshotter because Firecracker will not take an overlayfs rootfs.
The quantitative claims are Google's. The GKE GA post of 21 May 2026 reports more than 16x growth in sandboxes on GKE in under five months, a ratio with no base disclosed, and 300 sandbox allocations per second per cluster with 90 percent completing in 200 ms, naming LangChain and Lovable as customers. The project roadmap lists the controller work behind the 300 per second figure as done, warm-pool rolling updates as in progress, and a 50 ms claim latency, scale-to-zero, a sandbox router and an MCP server as planned. The same May post introduced Agent Substrate, a thinner control plane for "millions of sub-second tool calls" that reuses the sandbox runtime and snapshotting but, in Google's words, bypasses some Kubernetes limitations. I read that as a disclosure: the API server is the bottleneck at the scale where every tool call is a sandbox.
The YAML
This is a working v1beta1 template, pool and claim. The template pins gVisor through a RuntimeClass, disallows claim-time environment injection so claims cannot fall off the warm path, and attaches a DRA claim for a GPU with at least 40 GiB. Install the controller first with kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/download/v1.0.5/manifest.yaml and the same path's extensions.yaml.
# agent-sandbox v1.0.5 (extensions.agents.x-k8s.io/v1beta1) on Kubernetes >= 1.34
# DRA attribute names below follow NVIDIA k8s-dra-driver-gpu v0.5.0; adjust for your driver.
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: gpu-40g
namespace: agents
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.nvidia.com
selectors:
- cel:
expression: "device.capacity['gpu.nvidia.com'].memory.compareTo(quantity('40Gi')) >= 0"
---
apiVersion: extensions.agents.x-k8s.io/v1beta1
kind: SandboxTemplate
metadata:
name: coder-gvisor
namespace: agents
spec:
envVarsInjectionPolicy: Disallowed # claims that add env are rejected, not cold-started
volumeClaimTemplatesPolicy: Disallowed
service: true # per-sandbox headless Service -> status.serviceFQDN
networkPolicyManagement: Managed # controller owns the sandbox NetworkPolicy
podTemplate:
spec:
runtimeClassName: gvisor # or kata-qemu / kata-fc; GPU needs a gVisor build with nvproxy
restartPolicy: OnFailure
securityContext:
runAsNonRoot: true
runAsUser: 10000
resourceClaims:
- name: gpu
resourceClaimTemplateName: gpu-40g
containers:
- name: runtime
image: ghcr.io/example/agent-runtime:1.4.2 # your image; see the sandboxd docs for exec/file APIs
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "4Gi" }
claims:
- name: gpu
---
apiVersion: extensions.agents.x-k8s.io/v1beta1
kind: SandboxWarmPool
metadata:
name: coder-pool
namespace: agents
spec:
replicas: 20 # paid-for idle capacity; size from your claim rate
sandboxTemplateRef:
name: coder-gvisor
---
apiVersion: extensions.agents.x-k8s.io/v1beta1
kind: SandboxClaim
metadata:
name: session-7f3a
namespace: agents
spec:
warmPoolRef:
name: coder-pool # only a pool ref: anything else forces the cold path
additionalPodMetadata:
labels:
sandbox.users.io/tenant: acme
Notice who owns what: the platform team owns everything in the template, runtime class, network policy and device class included, and the agent developer owns a five-line claim. That is the platform-engineering contract in one file. And the warm pool is twenty running pods doing nothing, which is where the "sub-second" comes from and what you pay for it.
Plain Pods, Agent Sandbox, or a microVM fleet
The only published density comparison is Google's own, in a 31 July 2026 post, on a single n2-standard-48 running OpenClaw agents: 61 agents on Kata microVMs before reliability degraded, 88 on Agent Sandbox with gVisor, 133 with warm pools, and 274 with suspend-and-resume, which is where the 3.5x density and 75 percent cost-per-agent headlines come from. The trade is stated in the same post: startup goes from under one second to under five. Treat it as a vendor benchmark with a Kata-on-the-same-nodes baseline, not a comparison against a tuned Firecracker fleet, and note that it sets a cost-optimised configuration against a default one.
| Option | Isolation boundary | Startup path | Agents per n2-standard-48 (Google, Jul 2026) | Source |
|---|---|---|---|---|
| Plain Pod, runc | Shared host kernel, seccomp, NetworkPolicy | New pod, about 1 s | not tested | kubernetes.io |
| Agent Sandbox + gVisor | User-space kernel per pod | Warm pool: 300/s per cluster, p90 200 ms | 88; 133 with warm pool; 274 with suspend/resume | GKE GA, GKE cost |
| Agent Sandbox + Kata (QEMU or Firecracker) | KVM guest per pod | kata-fc about 125 ms boot (repo claim), needs KVM nodes | 61 (Kata, Google's baseline) | example README |
| Hosted microVM (E2B, Vercel Sandbox, PandaStack) | Firecracker VM per sandbox, own kernel | Snapshot restore; Vercel example 1 vCPU 2 GB at 10% CPU: $0.0552/h vs $0.1704 wall-clock | n/a, billed per sandbox | Vercel docs, E2B infra |
Where the cluster wins is clear. If the GPUs, the network policy, the identity provider and the compliance evidence already live on the cluster, and the agents are the long-lived, stateful singletons the project describes, then a SandboxTemplate is the shortest path, and the 1.37 HPA scale-to-zero beta plus operatingMode: Suspended covers the idle hours. Where it loses is also clear. gVisor is a user-space kernel sharing the host's, Kata on QEMU is a VM per pod with the density you saw above, and the Firecracker path inside Kubernetes needs KVM nodes and a non-default snapshotter, so you end up operating a microVM fleet anyway, with the API server in the path of every claim. That is the architecture my isolation comparison warned about and the one Agent Substrate exists to route around. The hosted providers, E2B with its Apache-2.0 Firecracker runtime, Vercel Sandbox on Firecracker with Active CPU pricing, Cloudflare's Sandbox SDK on microVM containers still in public beta, and PandaStack, sell the kernel boundary, snapshot-restore and fork semantics, and no cluster to upgrade three times a year. I covered who needs which boundary in the provider roundup. Be careful with every startup number in this section: as I argued in benchmarks that mislead, a 200 ms warm-pool claim, a 125 ms Firecracker boot and a snapshot restore measure three different things.
Frameworks and gateways above the sandbox
A sandbox is where the agent's code runs; something else has to define the agent, give it a model and tools, and put a proxy in front of it. On the framework layer, kagent, a CNCF sandbox project since 22 May 2025 from Solo.io with 3,935 stars, defines an agent as Kubernetes resources: Agent, a shareable AgentTemplate, a Harness that says whether it runs on kagent's own Go or Python ADK, Codex, Claude or a custom runtime, ModelConfig for the provider and RemoteMCPServer for tools, with sessions that suspend, resume and fork from a checkpoint. It shipped v0.10.3 on 2 October while cutting 1.0.0 alphas almost daily, alpha7 on 1 October, so the 1.0 API is weeks away but not here. Dapr Agents 1.0 went GA at KubeCon Amsterdam on 23 March 2026, a Python framework on Dapr's durable workflow engine with state in any of 30-plus stores and SPIFFE identities, built with NVIDIA and demonstrated by ZEISS; PyPI has 1.0.7 from 2 October under Apache-2.0. The two solve different problems: kagent makes the agent a Kubernetes object, Dapr Agents makes the agent's run survive a pod restart.
On the gateway layer the interesting event is a split. agentgateway, a Linux Foundation project in Rust under Apache-2.0 with 5,172 stars, speaks MCP, A2A and the OpenAI, Anthropic, Bedrock and Gemini APIs and reached v1.6.0 on 2 October. kgateway, in the CNCF sandbox since 4 March 2025, used its v2.3.0 release on 15 May 2026 to remove its Inference Extension support, "which had already moved to agentgateway". The north-south ingress gateway and the agent gateway are now separate projects with separate release trains, and Solo's KubeCon Europe recap adds a third, agentregistry, donated to the CNCF in March. On the inference side the Gateway API Inference Extension made InferencePool v1 in v1.0.0 on 9 September 2025 with conformance reports from GKE, Istio, kgateway, Envoy AI Gateway and agentgateway, and then in v1.6.0 on 17 August 2026 moved its Endpoint Picker, body-based routing and latency predictor into llm-d, keeping only the spec, CRDs and conformance. llm-d, founded by Red Hat, Google, IBM, CoreWeave and NVIDIA, entered the CNCF sandbox on 12 March 2026 and shipped v0.10.0 on 29 September; KServe, incubating, integrated llm-d v0.6 and vLLM 0.19 in v0.18 on 29 April and is at v0.21.0 as of 25 September. The serving stack I described in LLM Inference in 2026 is now three projects with clean seams: the gateway routes, llm-d schedules across prefill and decode, KServe owns the model lifecycle.
The 80 percent of platform teams nobody counted
Every platform-engineering deck carries the Gartner line that "by 2026, 80% of large software engineering organizations will establish platform engineering teams, up from 45% in 2022". Gartner's own page on the trend blocked my fetch, so the wording above is from DevOpsDigest's May 2024 reproduction of the Gartner release. It is 2026, and I could not find anyone, Gartner included, who has published a measurement of whether it came true. The nearest public data is the CNCF survey's finding that 47 percent of respondents now rank "cultural changes with the development team" as their top obstacle, ahead of training, security and complexity, which is a platform-team problem described without the phrase.
Backstage is the usual proxy for adoption and it also disappoints the deck-makers. It has been a CNCF incubating project since 15 March 2022, not graduated; the ADOPTERS.md file listed about 292 organisations when I counted on 5 October; and the CNCF's documentary announcement in March ranks it sixth by velocity among more than 230 projects in 2025. The "3,400 organisations and two million developers" figures that circulate come from vendor pages and I could not trace them to a primary source, so I am not repeating them as fact.
What the agent work changes about platform engineering is not the headcount but the tenant. An agent is a user who writes code at machine speed and runs it immediately, and the Agent Sandbox examples are candid about what that means: in the repository's multi-user agent demo the NetworkPolicy that restricts ingress to the gateway "is the per-user isolation boundary" when credentials are shared inside the sandbox, and the template's injection policies exist so that a claim cannot quietly widen what the platform decided. RuntimeClass, network policy, device class, execution-scoped tokens, warm-pool sizing and the suspend policy are platform decisions; the agent team gets a claim. That is the golden-path and self-service split platform engineering argued for before the Gartner number, and it is why the DevOps discipline underneath still matters: templates in Git, pinned versions (the Kubernetes blog's quickstart says to pin a tag), rolling updates of warm pools, which the roadmap still marks in progress, and a cluster upgraded three times a year, where sandbox CRDs, DRA driver and gateway all move together.
What I take from the year is that the question changed. Kubernetes is now a credible control plane for agents that are long-lived, stateful and few per node, and the project that made it so went from a KubeCon preview to v1.0 in nine months. It is not yet a credible substrate for the other shape of agent, the one that makes a thousand short tool calls a minute and needs a fresh kernel for each, and Google saying so by shipping Agent Substrate is more convincing than anything I could write. This quarter I would do three things if I ran a platform team: upgrade past 1.34 before its 27 October end-of-life so DRA is not a feature gate anywhere, stand up agent-sandbox v1.0.x on a gVisor node pool with a twenty-pod warm pool and measure the claim latency and the idle bill myself rather than quoting Google's, and keep the microVM budget line for the agents whose code nobody on the team wrote. The platform team's job did not change. The tenant did.
Related: Container vs microVM vs gVisor: Choosing Agent Isolation, Sandbox Creation Time Benchmarks Measure Different Things and DevOps Still Matters in the AI Era.
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
Infrastructure as Code in 2026: Two Forks at 1.16 and 1.13, a $6.4B Owner, 912 Public State Files, and One Agent That Ran terraform destroy
Infrastructure as code three years after the BSL relicence: Terraform 1.16.5 under IBM versus OpenTofu 1.13.1 under the Linux Foundation and who shipped what first, HCP Terraform at $0.10 to $0.99 per resource with the legacy free plan gone, CDKTF and System Initiative archived, Pulumi 3.267 and Crossplane 2.4, a Terraform MCP server that grew from registry lookups to workspace administration in 15 months, 44 percent running AI for infrastructure but 34 percent trusting it, 912 exposed state files with 41 live AWS keys, and the agent that ran terraform destroy on 2.5 years of production.
13 minOct 5, 2026MCP in 2026: 475M SDK Downloads a Month, 39,492 Servers in a Registry Still in Preview, and 12 Gateways Selling the Same 3 Features
The MCP ecosystem ten months after the Linux Foundation took it over, measured from primary sources: 475 million SDK downloads a month across npm and PyPI, three spec revisions ending in the stateless 2026-07-28 rewrite, an official registry I paged to 39,492 servers (24,242 remote) that is still labelled preview while Glama lists 96,340, every major client with different controls, twelve gateways from $0 to $0.005 per thousand calls, the tool-overload numbers (55k tokens before the first prompt, 85 percent recoverable), and a stateless Python server on mcp 2.3.0 that I ran and tested.
14 minOct 5, 2026DevOps Still Matters in 2026: AI Cut Delivery Stability 7.2%, Then Doubled Merged PRs and Added 91% to Review Time
Why DevOps is the big thing of the agent era: DORA 2024 found a 25% rise in AI adoption cost 1.5% throughput and 7.2% stability, DORA 2025 saw throughput turn positive while instability stayed up across nearly 5,000 respondents, and DORA's 2026 ROI model budgets a 15% three-month dip and a change failure rate rising from 5% to 6%; GitHub merged 518.7M PRs (+29%) and over 1M agent PRs in five months, Faros telemetry on 10,000 developers shows 98% more PRs and 91% longer reviews, METR found experienced developers 19% slower, and the Replit postmortem's fixes are 2015 DevOps controls.
13 min