Coding Agents in 2026: Cursor at $2B Then Sold to SpaceX, Cognition at $1B, 58% on Terminal-Bench, and the Only RCT Still Says Slower

Oct 5, 2026 · 14 min · Ajay Kumar

Most of the processes that boot a microVM on PandaStack, the Firecracker platform I build and run, were started by a coding agent that wanted somewhere to run a test suite it had just rewritten. I disclose the affiliation because it colours how I read this market: I sell to it and I spend my own budget in it. What I wanted to pin down was simple to state and hard to source. Who is winning in money rather than demos. What a seat costs now that nobody sells unlimited anything. Whether the agents have measurably got better, and whether the people using them have measurably got faster. Everything below comes from company posts, leaderboards I pulled this morning, and the few studies that measured, with the unverifiable claims named as such.

The scoreboard, as the companies report it

The revenue story has two protagonists and one surprising ending. Cursor disclosed over $500 million in ARR with its $900 million Series C at $9.9 billion on 6 June 2025, then crossed $1 billion with a $2.3 billion Series D at $29.3 billion on 13 November 2025. Its press page lists a Bloomberg report from 2 March 2026 headlined "Cursor Recurring Revenue Doubles in Three Months to $2 Billion", which I treat as a press figure, not a disclosure. Then on 14 August 2026 Cursor announced it had been acquired by SpaceX; no price was given, the post promises "more capable models at lower cost" from "the largest fleet of GPUs in the world", and the blog has since shipped Grok 4.6 and 4.7 as in-house models.

Cognition is the other curve. Devin's ARR went from $1 million in September 2024 to $73 million in June 2025; the Windsurf acquisition on 14 July 2025 added $82 million of ARR and 350 enterprise customers, three days after OpenAI's $3 billion deal for Windsurf collapsed and Google paid $2.4 billion to license the technology and hire its founders. Cognition raised $400 million at $10.2 billion that September, raised over $2 billion at $48 billion on 8 September 2026 with run-rate up from $492 million in May to almost $900 million, and crossed $1 billion on 25 September. Anthropic's one hard number for Claude Code is in its Series F post of 2 September 2025: over $500 million in run-rate revenue, usage up more than 10x in three months. I could not find a newer Claude Code figure on an Anthropic page, so the larger 2026 numbers in circulation are not here.

Reported annualised revenue by coding-agent vendor, with date ($M) company disclosure Bloomberg report 05001,0001,5002,000 Cursor, Jun 2025 (Series C) $500M+ ARR Cursor, Nov 2025 (Series D) $1B+ annualised Cursor, Mar 2026 (Bloomberg) $2B, doubled in three months Claude Code, Sep 2025 (Series F) $500M+ run-rate, usage 10x in 3 months Cognition, May 2026 $492M run-rate Cognition, Sep 2026 $1B run-rate (Sep 25); ~$900M on Sep 8 Lovable, Jul 2025 $100M ARR Windsurf at acquisition, Jul 2025 $82M ARR, 350+ enterprise customers Devin alone, Jun 2025 $73M ARR, from $1M in Sep 2024 Not plotted:GitHub Copilot and OpenAI Codex disclose no revenue; Microsoft reports nearly 140,000 Copilot organisations. Replit's widely quoted ARR could not be verified from a primary source. Lovable disclosed no ARR with its 2026 round.
Annualised revenue as stated in each vendor's own funding or milestone post, except Cursor's March 2026 figure, which is a Bloomberg report linked from Cursor's press page. Run-rate and ARR are used as each company uses them.

The incumbents report usage, not money. GitHub's Octoverse for 2025 counts over 180 million developers, nearly 80 percent of new ones using Copilot in their first week, and over a million Copilot coding-agent pull requests between May and September 2025 against 43.2 million merged per month overall. Microsoft's April 2026 earnings remarks put nearly 140,000 organisations on Copilot with enterprise subscribers nearly tripled year over year. Google said on its February 2026 call that Antigravity, the agent-first IDE launched on 18 November 2025 with Gemini 3, Claude Sonnet 4.5 and GPT-OSS selectable, had passed 1.5 million weekly active users in just over two months. Lovable, covered in the no-code post, went from $100 million ARR in July 2025 to $6.6 billion in December and $13.3 billion in August 2026, disclosing no ARR with the latest round. OpenAI's Codex usage statistics live on pages my fetcher could not reach, so they are absent rather than guessed.

Vendor Latest reported revenue Valuation or owner Entry price Source
Cursor $1B+ annualised (Nov 2025); $2B (Bloomberg, Mar 2026) Acquired by SpaceX, Aug 2026 Individual $20, Teams $40 per user Series D, SpaceX, pricing
Claude Code $500M+ run-rate (Sep 2025) Anthropic, $183B post Series F Pro $20, Max $100 or $200, Team $20 to $25 per seat Series F, pricing
Cognition (Devin, Windsurf) $1B run-rate (Sep 2026) $48B, Series E Sep 2026 Not published on the posts cited Series E, milestone
GitHub Copilot Not disclosed; ~140,000 organisations Microsoft Pro $10, Pro+ $39, Max $100, Business $19, Enterprise $39 per user Microsoft Q3 FY26, plans
OpenAI Codex Not disclosed OpenAI Included in ChatGPT Free through Enterprise; Pro has no five-hour limit Codex pricing docs
Google Antigravity, Gemini CLI, Jules Not disclosed; 1.5M weekly users (Feb 2026) Alphabet Free for individuals; Business from $30 per seat Q4 2025 call, pricing
Lovable $100M ARR (Jul 2025); none disclosed since $13.3B, Series C Aug 2026 Not covered here Series C
Amazon Kiro Not disclosed Amazon Free 50 credits; Pro $20 for 1,000; Power $200 for 10,000; $0.04 per extra credit kiro.dev
JetBrains Junie Not disclosed JetBrains Lite free; AI Pro $8.33; AI Ultimate $25 per user, BYOK junie.jetbrains.com
ByteDance Trae Not disclosed ByteDance Pro $20 for $20 of usage; Pro+ $60; Ultra $200 trae.ai

Four re-pricings in fifteen months

Every price in that table is a different shape from the one sold in early 2025, the transition I described for SaaS generally in the repricing post: seats to metered tokens, with a credit layer to make the meter palatable.

Cursor went first. On 16 June 2025 it introduced a $200 Ultra tier with "20x more usage than Pro" and redefined Pro as "at least $20 of model inference at API prices per month" instead of 500 requests; on 4 July it apologised for a rollout that "was not communicated clearly" and refunded three weeks of surprise charges. GitHub's premium-request billing began on 18 June 2025: 300 requests a month on the $10 Pro plan, 1,500 on the $39 Pro+, $0.04 each beyond. Anthropic's weekly caps took effect on 28 August 2025; by TechCrunch's account of the notice, 40 to 80 hours of Sonnet 4 a week on the $20 plan, 140 to 280 hours plus 15 to 35 of Opus on the $100 Max, 240 to 480 plus 24 to 40 on the $200 Max, aimed at the "less than 5 percent" of subscribers running Claude Code "continuously in the background, 24/7" or reselling access. The current pricing page keeps it general, "weekly and monthly caps ... at our discretion", with usage credits at API rates beyond.

The fourth and cleanest re-pricing is GitHub's. On 27 April 2026 it announced that from 1 June premium requests would become "GitHub AI Credits" computed from input, output and cached tokens at listed API rates, because "a quick chat question and a multi-hour autonomous coding session can cost the user the same amount". Headline prices did not move: Pro $10 with $10 of credits, Pro+ $39 with $39, Business $19, Enterprise $39. On 12 May it added "flex allotments", $5 on Pro, $31 on Pro+ and $100 on a new $100 Max, "designed to adapt as the economics of AI evolve", which is an allowance the vendor can shrink without touching the price. Amazon's Kiro and ByteDance's Trae launched straight onto credits. Only OpenAI differs: the Codex pricing docs include Codex in every ChatGPT tier down to Free, quote Plus as 350 to 3,000 messages per five hours on GPT-6 Luna against 15 to 160 on GPT-6.1 Sol, and say Pro "currently" has no five-hour limit. Google gives Antigravity away to individuals under five-hour limits, rations Jules at 15, 100 or 300 tasks a day by tier, and has kept Gemini CLI's 1,000 free requests a day since June 2025. The flat $20 was a customer-acquisition subsidy; the credits are the invoice arriving.

What the benchmarks say, and what the leaderboards stopped saying

The public leaderboards and the vendor announcements have drifted apart. I pulled the official SWE-bench site this morning: the highest open submission is 79.2 percent, live-SWE-agent with Claude Opus 4.5 on 15 December 2025, and the newest entry of any kind is dated 19 February 2026. Nothing from the 2026 model generation is on it. Anthropic reported 81.42 percent for Opus 4.6 in February "with a prompt modification" and Mistral 72.2 percent for the 123-billion-parameter Devstral 2. The benchmark is saturated, the labs have stopped submitting, and the number in a launch post is a number the lab ran itself.

Two harder benchmarks have taken over. Scale's SWE-bench Pro draws 1,865 tasks from 41 professional repositories, 731 public; its top public scores are 61.5 percent for Muse Spark 1.1, 59.1 for GPT-5.4 at highest effort, 51.9 for Claude Opus 4.6 and 46.1 for Gemini 3.1 Pro. Terminal-Bench 4.0, from Stanford, Harbor and the Laude Institute, measures long tasks in a real shell, the thing agents are sold to do. Its data this morning puts Codex with GPT-6 Astra at maximum effort on top at 58.2 percent plus or minus 2.8 over 330 trials costing $3,267, with Claude Code and Fable 5.1 at 57.9 plus or minus 3.8 for $6,243, statistically tied at nearly twice the cost. Versioning matters as much as the score: Cognition's SWE-2 post reports 92.8 percent on Terminal-Bench 2.1 and 27.3 percent on Terminal-Bench 4 for the same model. Anthropic's Opus 5.5 announcement on 22 September claims 66.4 percent on Terminal-Bench 4.0 at $4 and $20 per million tokens; it was not on the public board when I pulled it, so it is a vendor number for now.

Benchmark Top score when pulled Agent and model Note Source
Terminal-Bench 4.0 58.2% ± 2.8 Codex, GPT-6 Astra, max effort, 3 Sep 2026 $3,267 over 330 trials tbench.ai
Terminal-Bench 4.0 57.9% ± 3.8 Claude Code, Fable 5.1, max effort, 1 Sep 2026 $6,243 over 330 trials tbench.ai
Terminal-Bench 4.0 66.4% (vendor) Claude Opus 5.5, 22 Sep 2026 Not yet on leaderboard Anthropic
SWE-bench Pro (public) 61.5% Muse Spark 1.1 731 public tasks Scale
SWE-bench Pro (public) 51.9% Claude Opus 4.6 (thinking) Scale
SWE-bench Verified (official board) 79.2% live-SWE-agent, Claude Opus 4.5, 15 Dec 2025 Newest entry Feb 2026 swebench.com
SWE-bench Verified (vendor) 81.42% Claude Opus 4.6, "with a prompt modification" Anthropic

So the honest ceiling for an unattended agent on realistic terminal work is a little under three tasks in five, at a few thousand dollars a run, with the two frontier harnesses statistically tied.

The productivity evidence: one RCT, redesigned

The only randomised trial of agentic tools on real work is still METR's. The July 2025 study gave 16 experienced maintainers 246 issues from their own repositories, two hours each on average, with Cursor Pro and Claude 3.5 and 3.7 Sonnet, and found they took 19 percent longer with AI allowed, interval plus 2 to plus 39. They had predicted a 24 percent speedup and afterwards believed they had got 20. In February 2026 METR reported the second wave on late-2025 tools, 57 developers, 143 repositories, over 800 tasks: the original cohort at minus 18 percent, interval minus 38 to plus 9, newly recruited developers at minus 4, interval minus 15 to plus 9. Then it announced a redesign, because the experiment had started to break: developers refused to join if it meant working without AI, 30 to 50 percent admitted withholding tasks they did not want to do by hand, and people using agents could not say how long a task took because they were doing something else while it ran. METR's reading is that the data is "only very weak evidence" for a productivity increase, and the point estimates are still negative.

Against that sits self-report, which METR also measured. Its May 2026 survey of 349 technical workers found a median self-reported 1.4 to 2x change in the value of their work and a 3x speed change, with the reminder that in its own RCT people overestimated AI's effect on their time by 40 percentage points. DORA 2025, from nearly 5,000 respondents, found 90 percent using AI at work, over 80 percent believing it made them more productive, 30 percent with little or no trust in the code, and a positive association between AI adoption and throughput alongside a persistent negative one with delivery stability; the 2026 report was not on dora.dev when I checked. Stack Overflow's 2025 survey, the latest published, has 84 percent using or planning to use AI tools, yet 46 percent distrust the accuracy of the output against 33 percent who trust it, 66 percent name "almost right, but not quite" as their biggest frustration, and 14.1 percent used agents daily. The gap between what developers feel and what a stopwatch says is the most consistent finding in the field, and the revenue is priced on the feeling.

Security and the maintainer tax

The code-quality evidence is less ambiguous. Veracode's 2025 GenAI Code Security Report ran over 100 models through tasks in Java, Python, C# and JavaScript and found 45 percent of samples introduced an OWASP Top 10 weakness, 72 percent in Java, with cross-site scripting defended against in only 14 percent of relevant cases, and newer, larger models "were no better at writing secure code". CodeRabbit's December 2025 report compared 320 AI-co-authored pull requests with 150 human-only ones and found 10.83 issues per AI PR against 6.45, about 1.7x, with logic errors 75 percent more common, security findings up to 2.74x and readability problems over 3x. Both studies have limits, a review vendor grading what it sells a fix for, a lab task rather than a merged diff, but their direction agrees with the "almost right" number.

The cost lands on reviewers. Daniel Stenberg ended the curl bug bounty on 31 January 2026 after paying over $100,000 for 87 confirmed vulnerabilities since April 2019, because the share of real reports fell from above 15 percent to below 5 percent in 2025 under what he calls an explosion of AI slop; in July 2026 the maintainers took the entire month off from vulnerability reporting. Hacktoberfest's 2026 rules say pull requests "will no longer count toward Hacktoberfest rewards" because "it's easier than ever to submit low-effort spam PRs". GitHub's own Copilot app post in June concedes "too much time spent reviewing agent-generated code" alongside 1.4 billion commits a month. Agents have made the review queue the scarce resource.

The open-source layer and the Chinese entrants

Underneath the paid products is a bring-your-own-key market that shows up in GitHub stars rather than ARR. I pulled the counts on 5 October 2026 with the GitHub API: OpenCode, MIT, 211,780; Anthropic's Claude Code repository 149,442 and no licence file; OpenAI's Codex CLI, Apache-2.0, 127,877; Gemini CLI 107,238; Zed 91,296; OpenHands 90,010; Cline 69,866; goose 54,949; Aider 49,382, last pushed in May 2026. The convention that lets them interoperate, AGENTS.md, is stewarded by the Agentic AI Foundation under the Linux Foundation and used by over 60,000 open-source projects. Zed raised a $32 million Series B from Sequoia in August 2025 and remains open source with a paid tier.

The Chinese entrants are real and derivative. Qwen Code, Apache-2.0 with 28,313 stars, states it "was originally based on Google Gemini CLI v0.8.2"; Moonshot's Kimi CLI, also Apache-2.0, has 11,433. ByteDance's Trae has the priced product, $20, $60 and $200 tiers of usage, and its agent with Doubao-Seed-Code holds 78.8 percent on the official SWE-bench board. Mistral's Vibe CLI, Apache-2.0, ships with Devstral 2 under a modified MIT licence at $0.40 and $2.00 per million tokens, a tenth of Opus 5.5's $4 and $20, and the 24-billion Devstral Small 2 under Apache-2.0 at 68 percent on Verified; the licensing side is in the open-weights post.

Counting cost per merged pull request

Every tool now bills in tokens, every harness logs its token counts, and almost nobody divides one by the other. The number that matters is not cost per request or per session but cost per pull request that actually merged, with abandoned sessions charged to the ones that landed, because that is how the money leaves. The script below does that from a JSONL of sessions. I ran it in my scratchpad against a six-session fixture I made up to exercise the code; the output is the real output of that run, not a measurement of anything in production. Prices are the list prices in the Opus 5.5 and Devstral 2 announcements linked above.

#!/usr/bin/env python3
"""pr_cost.py: cost per merged pull request from agent session logs.

Input: JSONL, one agent session per line, with the token counts your harness
already emits (Claude Code, Codex and OpenCode all log these) and the PR the
session worked on.  Prices are list API prices per million tokens as published
on 2026-09-22 (Opus 5.5) and 2025-12-09 (Devstral 2); update them, do not trust them.
"""
import json, sys, collections

PRICE = {  # USD per 1M tokens: (uncached input, cached input read, output)
    "claude-opus-5-5": (4.00, 0.20, 20.00),   # anthropic.com/news/claude-opus-5-5
    "devstral-2":      (0.40, 0.40, 2.00),    # mistral.ai, no cache discount listed
}

def session_cost(s):
    p_in, p_cache, p_out = PRICE[s["model"]]
    cached = s.get("cached_input_tokens", 0)
    fresh = s["input_tokens"] - cached
    return (fresh * p_in + cached * p_cache + s["output_tokens"] * p_out) / 1e6

by_pr = collections.defaultdict(lambda: {"cost": 0.0, "sessions": 0, "merged": False})
orphan = 0.0                                   # sessions that never produced a PR
for line in open(sys.argv[1]):
    s = json.loads(line)
    c = session_cost(s)
    if s["pr"] is None:
        orphan += c; continue
    row = by_pr[s["pr"]]
    row["cost"] += c; row["sessions"] += 1; row["merged"] |= s["merged"]

total = sum(r["cost"] for r in by_pr.values()) + orphan
merged = [r for r in by_pr.values() if r["merged"]]
for pr, r in sorted(by_pr.items()):
    print(f"{pr:14} {r['sessions']} session(s) ${r['cost']:7.2f} {'merged' if r['merged'] else 'abandoned'}")
print(f"total spend ${total:.2f} incl. ${orphan:.2f} with no PR; "
      f"{len(merged)}/{len(by_pr)} PRs merged; "
      f"cost per merged PR ${total/len(merged):.2f}")   # all spend, not just the winners
$ python3 pr_cost.py sessions.jsonl
org/api#412    1 session(s) $   2.34 merged
org/api#413    2 session(s) $   4.16 merged
org/web#77     1 session(s) $   1.42 merged
org/web#78     1 session(s) $   2.36 abandoned
total spend $10.83 incl. $0.55 with no PR; 3/4 PRs merged; cost per merged PR $3.61

Even with invented numbers the fixture shows two things: the 20x gap between Opus 5.5's $4 uncached and $0.20 cached input price decides more of the bill than the output price does, so a harness that breaks the cache is a cost bug; and the abandoned pull request is a third of the spend, which on work where frontier agents fail two tasks in five makes the denominator the number to watch.

What I take from the primary sources is less dramatic than the valuations. The money is real and concentrated: two companies at or past a billion in run-rate, one now inside SpaceX. Pricing has converged on metered tokens everywhere except OpenAI's bundle, because the flat seat never covered the heavy users. The agents top out a little under 60 percent on the hardest public terminal benchmark, and the only controlled study still cannot find a speedup while developers report two to three times; those facts coexist because review, not generation, is the bottleneck, which is why Stenberg closed his bounty and Hacktoberfest stopped counting pull requests. This quarter I am doing the boring version of that lesson on my own repositories: cost per merged PR in the dashboard next to CI minutes, cache-hit rate as a first-class metric, and no agent-opened pull request merges without a human who can explain the diff. The platform discipline the agents were supposed to make unnecessary is what makes their output usable.

Related: DevOps Still Matters in the AI Era, CI/CD When Agents Open the Pull Requests and Seats Are Dying: The 2026 SaaS Repricing From the Compute Layer.

I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.

Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.


Related

Oct 5, 2026

The 2026 SaaS Selloff: IGV Fell 36%, Then Rose 45%. Revenue Grew 11% to 37% the Whole Time

The SaaSpocalypse checked against filings: Anthropic's Cowork plugins on 30 January, Thomson Reuters' worst day ever on 3 February, $285 billion gone in a day and nearly a trillion in a week, a second leg on Mythos in April, IGV down 36 percent to its 10 April low and up 45 percent since; against that, Q2 2026 revenue growth of 11 percent at Salesforce, 24 at ServiceNow, 36 at Datadog and 37 at Snowflake, median retention flat at 110 percent, median multiples from 4.9x to 3.0x to 4.2x, 4,000 Salesforce support roles, Klarna's reversal and $65 billion of Anthropic run-rate.

13 min
Oct 5, 2026

DevOps Still Matters in 2026: AI Cut Delivery Stability 7.2%, Then Doubled Merged PRs and Added 91% to Review Time

Why DevOps is the big thing of the agent era: DORA 2024 found a 25% rise in AI adoption cost 1.5% throughput and 7.2% stability, DORA 2025 saw throughput turn positive while instability stayed up across nearly 5,000 respondents, and DORA's 2026 ROI model budgets a 15% three-month dip and a change failure rate rising from 5% to 6%; GitHub merged 518.7M PRs (+29%) and over 1M agent PRs in five months, Faros telemetry on 10,000 developers shows 98% more PRs and 91% longer reviews, METR found experienced developers 19% slower, and the Replit postmortem's fixes are 2015 DevOps controls.

13 min
Oct 5, 2026

Infrastructure as Code in 2026: Two Forks at 1.16 and 1.13, a $6.4B Owner, 912 Public State Files, and One Agent That Ran terraform destroy

Infrastructure as code three years after the BSL relicence: Terraform 1.16.5 under IBM versus OpenTofu 1.13.1 under the Linux Foundation and who shipped what first, HCP Terraform at $0.10 to $0.99 per resource with the legacy free plan gone, CDKTF and System Initiative archived, Pulumi 3.267 and Crossplane 2.4, a Terraform MCP server that grew from registry lookups to workspace administration in 15 months, 44 percent running AI for infrastructure but 34 percent trusting it, 912 exposed state files with 41 live AWS keys, and the agent that ran terraform destroy on 2.5 years of production.

13 min