FinOps for AI in 2026: 98% Manage Token Spend, 3 in 4 Can't Prove the Value, GPUs Run at 5%, and FOCUS Gets a Token Column in 1.5
PandaStack, the Firecracker microVM cloud I build and operate, bills by the second of sandbox time, and I have seen enough customer workloads to know that on an agent job my line is the small one. (Disclosure: I founded PandaStack and run it alone.) The tokens an agent burns inside my sandbox cost a multiple of the sandbox, and they arrive on a different invoice, from a different vendor, in a different unit, with none of the tags that fourteen years of cloud cost discipline taught us to demand. So this is the ledger for October 2026: who manages AI spend according to the FinOps Foundation's surveys, what a million tokens costs and what caching and batching do to it, why an agent loop costs what it does, what FOCUS can carry yet, which gateways refuse a request when a budget is gone, what GPUs are measured at, and the two unit-economics numbers I would put on a dashboard this quarter.
98 percent manage it, three in four cannot prove it
The FinOps Foundation's State of FinOps 2026, its sixth annual survey, covers 1,192 practitioners stewarding more than $83 billion of annual cloud spend. 98 percent of them now manage AI spend, against 63 percent in the 2025 report (861 respondents, $69 billion) and 31 percent in 2024. AI cost management is the top skill respondents say they lack and "FinOps for AI" the top forward-looking priority. The scope has widened with it: 90 percent manage or plan to manage SaaS, 57 percent private cloud, 48 percent data centre, 28 percent labour. Workload optimisation is still the leading priority, with diminishing returns reported, which is what six years of rightsizing the same instances gets you. On FOCUS, 57 percent of 2025 respondents planned adoption.
The uncomfortable number comes from the Foundation's partner on token economics. The Tokenomics Foundation's State of Tokenomics for September 2026, 472 responses across 11 industries, finds that three in four enterprises cannot confidently prove AI business outcomes to their CFO. 86 percent are evaluating or using model routers, and those who do are four times as likely to show value: the ones who can see per-request cost can also see per-request value. Asked what they want from vendors, 23 percent said granular data; 4 percent said discounts. Two widely repeated figures, that 73 percent of enterprises blew their AI budgets and that 80 to 90 percent of AI spend is inference, I could not find in the Foundation's report and have left out.
The price list, October 2026
I checked each vendor's pricing page on 5 October 2026. Prices are per million tokens, standard tier; the cache column is the price of reading a cached prefix, which is the number that decides an agent's bill.
| Model | Input | Cached input | Output | Batch | Checked | Source |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $0.25 (0.025x) | $50 | 50% off | 2026-10-05 | Anthropic |
| Claude Opus 5.5 | $4 | $0.20 (0.05x) | $20 | 50% off | 2026-10-05 | Anthropic |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 | 50% off | 2026-10-05 | Anthropic |
| Claude Haiku 4.5 | $1 | $0.10 | $5 | 50% off | 2026-10-05 | Anthropic |
| gpt-6-astra | $10 | $1.00 | $50 | 50% off | 2026-10-05 | OpenAI |
| gpt-5.6-sol | $4 | $0.40 | $20 | 50% off | 2026-10-05 | OpenAI (promo to 21 Nov 2026) |
| gpt-5.4 / gpt-5.4-mini | $2.50 / $0.75 | $0.25 / $0.075 | $15 / $4.50 | 50% off | 2026-10-05 | OpenAI |
| Gemini 3.1 Pro Preview (≤200k) | $2 | $0.20 + $4.50/M/hour storage | $12 | 50% off | 2026-10-05 | |
| Gemini 3.8 Flash | $0.75 | $0.075 + $0.50/M/hour | $3.75 | 50% off | 2026-10-05 | Google (doubles 1 Jan 2027) |
| DeepSeek V4-Pro (off-peak / peak) | $0.66 / $1.32 | $0.022 / $0.044 | $1.98 / $3.96 | none | 2026-10-05 | DeepSeek |
| DeepSeek V4.1-Flash (off-peak / peak) | $0.15 / $0.30 | $0.003 / $0.006 | $0.60 / $1.20 | none | 2026-10-05 | DeepSeek |
The structure matters more than the headline numbers. Output costs five times input everywhere except DeepSeek. A cache hit costs a tenth of input at most vendors, a twentieth on Opus 5.5 and a fortieth on Fable 5.1, but Anthropic charges 1.25 times input to write a five-minute cache entry and twice for an hour, and Google bills storage by the token-hour, so a cache nobody reads is a cost. DeepSeek halves everything outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, which makes a batch job's schedule a pricing decision. Anthropic's docs say batch and cache multipliers stack and the full 1M context is billed at the standard rate; OpenAI doubles the price above 272k input tokens and Google above 200k. Then the small print: Anthropic's post-4.7 tokenizer produces about 30 percent more tokens for the same text, two of the prices in the table are promotional, and regional inference carries a 10 percent uplift at all three US vendors. A comparison that ignores cache ratio, output share, time of day and promo expiry is wrong by a factor of two before it starts.
Where the agent money goes
Agents are expensive for a structural reason, not a model reason: every step resends the entire context. I priced a deliberately ordinary loop with a twenty-line Python script in my scratchpad: 20 tool-calling steps, an 8,000-token system prompt and tool set, 3,000 tokens of tool output added per step, 600 output tokens per step, at the list prices above. Without caching the loop sends 730,000 input tokens to produce 12,000 output tokens, 61 to 1. With a prompt cache, 65,000 tokens are written once and 665,000 are read back at the cache rate.
| Loop variant (20 steps) | Input or cache write | Cache read | Output | Total |
|---|---|---|---|---|
| Sonnet 5.5, no cache | $1.460 | 0 | $0.120 | $1.580 |
| Sonnet 5.5, cached | $0.163 | $0.133 | $0.120 | $0.415 |
| Sonnet 5.5, cached + batch | $0.081 | $0.067 | $0.060 | $0.208 |
| Haiku 4.5, cached | $0.081 | $0.067 | $0.060 | $0.208 |
| Opus 5.5, cached | $0.325 | $0.133 | $0.240 | $0.698 |
| Fable 5.1, cached | $0.812 | $0.166 | $0.600 | $1.579 |
Three things fall out. Input dominates: 92 percent of the uncached bill is context being re-read, so a cache-busting system prompt (a timestamp, a request ID, an unsorted tool list) is the most expensive bug in an agent codebase. The tiers compress: Fable 5.1 cached costs what Sonnet 5.5 costs uncached, so "which model" and "did caching work" are decisions of the same size. And the batch discount only applies to work that can wait, which for an interactive agent is none of it.
The public traffic data points the same way. Anthropic's June 2026 Economic Index reports 54 percent of Claude Code conversations served by Opus against 10 percent of chat, higher-wage occupations running 1.53 times as many turns with 1.34 times the output per turn, and app building using more than three times the tokens of the median conversation. The AgentSysBench paper from August finds non-LLM components dominating latency in five of ten agentic applications, sandbox working sets peaking at 28 GB per session, sessions idle for minutes to hours between steps, and tool-result caching removing 35.2 percent of redundant search calls. Idle state is infrastructure cost that never appears as a token, the part of the bill I do see, and why scale to zero matters for agents.
The vendors who sold flat-rate agents learned this in public. Cursor moved Pro from 500 requests to "$20 of frontier model usage at API pricing" on 16 June 2025, about 225 Sonnet 4 requests, and its 4 July apology refunded the surprise bills in between. Anthropic announced weekly limits for Claude Pro and Max on 28 July 2025, affecting under 5 percent of subscribers and aimed at people running Claude Code around the clock; the per-plan hour ranges exist only in coverage such as Slashdot. Replit bills Agent work per completed checkpoint, priced by effort, with third-party model calls passed through at the provider's public rate. Lovable's credits cost 0.5 to 2 per build message, $15 per 50 on Pro. Every one of these is the same move, seat to meter, the subject of my SaaS repricing post.
FOCUS has a currency for tokens, but no column until 1.5
FOCUS is the specification that is supposed to make all of this one dataset. The changelog runs 1.0 in June 2024, 1.1 in November 2024, 1.2 in June 2025, 1.3 in December 2025, 1.4 in June 2026. FOCUS 1.2 brought SaaS and PaaS into the schema and made virtual currencies first-class values of PricingCurrency, naming Snowflake credits, Databricks DBUs and "OpenAI GPT tokens". FOCUS 1.3 added contract commitments, split cost allocation columns for shared resources such as Kubernetes pods, and data recency flags. FOCUS 1.4 added InvoiceDetail and BillingPeriod datasets and 47 columns for invoice reconciliation. None carries a model identity or a token count. That is scoped for 1.5, planned for December 2026 in the release plan, alongside a price-sheet dataset. Until then a token is a PricingUnit string and a quantity, and joining it to a model is your problem.
Exports lag the spec. The adopters page lists AWS, Azure, Google Cloud, Oracle, IBM Cloud, Nebius and Grafana Cloud at 1.2; MongoDB, Vercel, Snowflake and Databricks at 1.3; Alibaba, Tencent, Huawei, OVHcloud, Cloudflare and CoreWeave at 1.0. Microsoft's own schema page labels its newest export "1.2-preview". The most a hyperscaler gives you today is a 1.2 row in which an AI charge is a SKU with a quantity in tokens.
The cloud-native tooling fills part of the gap, with edges. Bedrock's application inference profiles, tag-enabled since November 2024, flow cost allocation tags to Cost Explorer and CUR, but each profile is bound to one model, so you need one per model per team, and the finest grain is per usage type per day. The newer Projects construct for the OpenAI-compatible bedrock-mantle endpoint decouples the tag from the model via an OpenAI-Project header, at the same daily grain; for per-prompt detail AWS points you at your own invocation logs. Azure Cost analysis exposes serverless models as paygo-inference-input-tokens and paygo-inference-output-tokens meters. And Claude bought through AWS Marketplace or Microsoft Foundry is billed in Claude Consumption Units at $0.01 each, one aggregated CCU line on the cloud bill, with the per-model, per-key breakdown in the Claude Console. The pattern: the cloud bill says how much, the vendor console says which model, and only something in the request path says which team, task or customer.
Budgets in the request path
That something is a gateway, and the budget feature is what I would select on. LiteLLM (MIT, Python) is the open-source default: virtual keys per team, max_budget in dollars, a budget_duration that resets it, per-tag budgets, and a 422 by default (429 if you ask) when a team is over.
# litellm config.yaml: three model aliases, hard monthly dollar budgets, spend persisted to Postgres.
# Keys from docs.litellm.ai/docs/proxy/team_budgets and /docs/proxy/config_settings, checked Oct 2026.
model_list:
- model_name: sonnet # alias callers use; the real model is hidden behind it
litellm_params:
model: anthropic/claude-sonnet-5-5
api_key: os.environ/ANTHROPIC_API_KEY
- model_name: haiku
litellm_params:
model: anthropic/claude-haiku-4-5
api_key: os.environ/ANTHROPIC_API_KEY
- model_name: flash
litellm_params:
model: gemini/gemini-3.8-flash
api_key: os.environ/GEMINI_API_KEY
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY # admin key for /team/new, /key/generate
database_url: os.environ/DATABASE_URL # spend logs and budgets live here
litellm_settings:
max_budget: 20000.0 # whole-proxy ceiling in USD; 0 disables
budget_duration: 30d # resets monthly
budget_exceeded_status_code: 429 # default is 422; 429 lets SDK retry logic back off
default_team_settings: # applied to teams created via JWT upsert
- team_id: default-settings
max_budget: 500.0
budget_duration: 30d
tag_budget_config: # budgets keyed on request metadata tags
env:staging:
max_budget: 200.0
budget_duration: 30d
success_callback: ["prometheus"] # spend metrics by team, key, model
# Create a team with a $1,500 monthly budget; keys generated for it inherit the cap.
curl -s -X POST http://localhost:4000/team/new \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' \
-d '{"team_alias":"support-agents","max_budget":1500.0,"budget_duration":"30d","models":["sonnet","haiku"]}'
The hosted alternatives converged on the same idea this year. Cloudflare's AI Gateway added spend limits on 5 June 2026: it prices each request from the model's published rate, accumulates spend per model, provider or custom metadata such as team over fixed or rolling windows, and blocks or routes to a fallback model; the gateway is free. Portkey's gateway is MIT with 13,100 stars, but its budget limits are an Enterprise and select-Pro feature of the hosted plane. Helicone is Apache-2.0, 6,200 stars, free to 10,000 requests a month then $79 and $799 tiers, with no per-key spend limits on its pricing page. OpenRouter is closed and hosted; its fees are 5.5 percent on card-funded credits, and BYOK is free to $25,000 a month then 5 percent. Downstream, Vantage's Anthropic integration reads the Admin API for cost by model, workspace, key and service tier, and CloudZero ingests OpenAI, Anthropic, Azure OpenAI and Bedrock into cost per customer and per feature; Finout claims the same.
Whatever you pick, insist that it emits the OpenTelemetry GenAI attributes: gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.usage.cache_read.input_tokens, gen_ai.usage.cache_creation.input_tokens and gen_ai.usage.reasoning.output_tokens, with gen_ai.response.model and gen_ai.conversation.id. The registry still marks all of them Development, as covered in my observability post. Cost per task is then a short join against your span store.
# cost_per_task.py: price finished spans by gen_ai.* attributes; one row per conversation.
# Prices per MTok checked 2026-10-05 (platform.claude.com/docs/en/about-claude/pricing).
PRICE = {"claude-sonnet-5-5": dict(inp=2.0, cw=2.5, cr=0.20, out=10.0),
"claude-haiku-4-5": dict(inp=1.0, cw=1.25, cr=0.10, out=5.0)}
def span_cost(a: dict) -> float:
p = PRICE[a["gen_ai.response.model"]]
cached = a.get("gen_ai.usage.cache_read.input_tokens", 0)
written = a.get("gen_ai.usage.cache_creation.input_tokens", 0)
uncached = a["gen_ai.usage.input_tokens"] - cached - written # input_tokens includes cached tokens
return (uncached*p["inp"] + written*p["cw"] + cached*p["cr"]
+ a["gen_ai.usage.output_tokens"]*p["out"]) / 1e6
def cost_per_task(spans):
tasks = {}
for a in spans: # a = span attribute dict
tasks[a["gen_ai.conversation.id"]] = tasks.get(a["gen_ai.conversation.id"], 0.0) + span_cost(a)
return tasks # join to your outcome table on the id
GPUs measured at 5 percent
If tokens are the bill you cannot tag, GPUs are the bill you cannot fill. Cast AI's 2026 Kubernetes optimization report, direct measurement across tens of thousands of production clusters on AWS, Azure and GCP for calendar 2025 with GPU data to April 2026, before any optimisation was enabled, puts average GPU utilisation at 5 percent. CPU averaged 8 percent, memory 20, and CPU overprovisioning rose from 40 to 69 percent in a year. It is one vendor's customer base and I would like a second source, but it matches what I hear from anyone running inference on Kubernetes: the accelerator is reserved, the pod is scheduled, and the silicon waits for a batch.
The allocation tools have caught up faster than the utilisation. Kubecost's 2.4 release in October 2024 moved to usage-based GPU allocation from NVIDIA's DCGM exporter, so a $100 GPU used at 50 percent allocates $50 to the container, and separates workload idle from infrastructure idle; it is now IBM Kubecost with a free tier to 250 cores. OpenCost, the CNCF project underneath, made AI cost tracking its top 2026 priority and shipped it in 1.121.0 on 5 August: it joins vLLM's vllm:prompt_tokens_total and vllm:generation_tokens_total counters to its GPU allocation engine and exposes llm_cost_per_million_tokens per model, split into allocation cost (everything the model's GPUs cost) and usage cost (active compute, corrected for KV-cache hits), with input and output priced separately because, as the inference post argued, prefill and decode are different machines. The proof of concept ran on 109 GPUs and 30 models with llm-d. It is the first open tool that gives a self-hosting team the per-million-token number the API vendors print.
Showback first, then cost per resolved ticket
The Foundation's FinOps for AI guidance says showback should precede chargeback: tag by project, environment, team and workload type, show each team its share, and charge only once the data is trusted. Its recommended units are cost per inference, per token and per API call, training cost per unit of quality, utilisation and ROI. I would collapse that to two numbers for an agent product: cost per task, which the join above produces and which you should watch as a p95 because the tail is where runaway loops live, and cost per outcome, the number the customer is implicitly paying for.
The gap between the two is the whole business. Intercom prices Fin at $0.99 per outcome, where an outcome is a resolution, a configured handoff or a disqualification. Anthropic's own docs give a worked support example of about 3,700 tokens per conversation on Haiku 4.5, or $37 per 10,000 tickets: $0.0037 of model cost per ticket, 0.4 percent of Fin's price. The rest is retrieval, the sandbox, the humans who take the handoffs, the tickets Fin did not resolve and did not bill, and margin. Sierra is widely reported to price per resolution too but publishes nothing. If your model cost per outcome is 0.4 percent of the price, the token bill is not your risk; unresolved tickets and idle GPUs are. If it is 40 percent, you have a routing problem, and the finding that router users are four times as likely to show value to the CFO stops being surprising.
What I take from this quarter's reading. Everyone manages AI spend now; almost nobody can tie it to a result, and the standard that would make that a SQL query is one release away. The prices are public and the discounts large, but they are discounts on a shape, and the shape of an agent loop, 61 tokens in for every one out, means the prompt cache is worth more than the model choice. The clouds give you a daily aggregate with a tag; only a gateway gives you the request, and only OpenTelemetry attributes give you the task. So before December I would put every model call behind a gateway with a hard dollar budget per team and a 429 on breach, emit the gen_ai usage attributes with a conversation ID, track cost per task as a p95 and cost per outcome as a ratio, read the GPU utilisation number before buying more GPUs, and get model identity into my own tags now, because FOCUS 1.5 will only standardise a column you can already fill. Tag it, meter it, show it back: the oldest discipline in platform engineering did not stop applying because the resource is a token.
Related: Seats Are Dying. What the 2026 SaaS Repricing Looks Like From the Compute Layer, Agent Observability in 2026: The Spans Are Standard, the Standard Isn't Stable and GPU Cloud and Neocloud Economics in 2026.
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
DevOps Still Matters in 2026: AI Cut Delivery Stability 7.2%, Then Doubled Merged PRs and Added 91% to Review Time
Why DevOps is the big thing of the agent era: DORA 2024 found a 25% rise in AI adoption cost 1.5% throughput and 7.2% stability, DORA 2025 saw throughput turn positive while instability stayed up across nearly 5,000 respondents, and DORA's 2026 ROI model budgets a 15% three-month dip and a change failure rate rising from 5% to 6%; GitHub merged 518.7M PRs (+29%) and over 1M agent PRs in five months, Faros telemetry on 10,000 developers shows 98% more PRs and 91% longer reviews, METR found experienced developers 19% slower, and the Replit postmortem's fixes are 2015 DevOps controls.
13 minOct 5, 2026Observability in 2026: OpenTelemetry Graduated, the Collector Is Still v0.162, eBPF Instrumentation Is v0.14, and Datadog Bills $1.12B a Quarter
The non-AI state of observability in 2026 from primary sources: OpenTelemetry graduated from CNCF in May with 12,000 contributors from 2,800 companies, yet the Collector distribution is v0.162.0 with mixed component stability, profiles are alpha, messaging conventions are still Development and OBI, the donated Beyla, is v0.14.0. Datadog grew 36 percent to $1.12B in Q2, cost is the top concern for 31 percent of 1,363 surveyed engineers, and the levers that move the bill (tail sampling, OTTL filtering, cardinality limits, columnar storage, Arrow transport) live in the pipeline you run yourself.
13 minSep 9, 2026AI Power in 2026: 485 TWh, a 0.3 GW Stargate, a 116 GW Turbine Backlog, and 1,800 MW That Dropped Off the Grid in Seconds
The electricity numbers behind AI, kept honest: IEA's 485 TWh for 2025 and 950 by 2030, LBNL's 4.7 percent of US power heading to 12, what a prompt actually costs (0.24 Wh at Google, 30 times more in reasoning mode), which gigawatt campuses are energised versus announced (Stargate Abilene at 0.3 of 1.2 GW), where the power comes from (a 116 GW gas-turbine backlog with 2031 slots, nuclear restarts in 2027, one SMR construction permit), what the grid operators are doing about 233 GW queues and 1,800 MW load-loss events, rack power from 132 kW to 600 kW, and a script to measure your own job's energy.
14 min