AIOps in 2026: 47% on ITBench, 10% on Hard On-Call, $6.50 an Investigation, and Four Outages Where the Automation Was the Incident

Oct 5, 2026 · 14 min · Ajay Kumar

I am the on-call rotation for PandaStack, the Firecracker microVM cloud I built and run alone, so I have a personal stake in the question this post asks: when the page arrives at 2am, what should a machine be allowed to do before I am awake? (Disclosure: PandaStack is mine.) Every observability vendor and a dozen startups now sell an "AI SRE" to answer that. I spent the week on the product pages, pricing tables, funding announcements, five hyperscaler postmortems and seven academic benchmarks, and the short version is this: agents are now useful investigators, the measured evidence for faster resolution is thinner than the claims, and the largest outages of the past year were all cases where automation did exactly what it was built to do. The argument is for agents that investigate at full speed and remediate only through a human-held gate, and for treating observability data quality as the ceiling on what they can find.

What the incumbents shipped, and what it costs

Datadog's Bits AI SRE went generally available on 2 December 2025 after being "tested against more than 2,000 customer environments" with "tens of thousands of investigations", per the press release. A monitor triggers it; it reads the runbook, forms competing hypotheses and marks each validated, invalidated or inconclusive. Datadog's own description is that "what once took more than 30 minutes of manual triage now happens automatically"; the remediation end, a proposed pull request that "engineers can review and merge", is in private preview. The pricing page is the most honest document in the set: AI Credits at $500 per 500 credits a month, $1.30 per credit on demand, about 6.5 credits per autonomous investigation, so roughly $6.50 each, no rollover. That is a price for investigation, not resolution.

PagerDuty took the other end of the incident. Its SRE Agent, "Paige", ingests events, runbooks and logs, recommends diagnostic and remediation steps and surfaces "nudges" as clickable buttons; each request or click costs 4 AI Actions, and the documentation is clear that it recommends and the human activates. The AI agents page lists four GA agents sold as credits with no dollar figures; the monthly prices circulating in comparison blogs I could not verify from PagerDuty.

Grafana Labs made six AI tools GA on 27 July 2026, including Assistant Investigations and a Cloud MCP server, in every Grafana Cloud tier, per the press release. Dynatrace announced domain-specific agents on 28 January with "policy-driven controls and approvals" (release), then on 27 July an Autonomous SRE Agent and a Cloud SRE Agent that "coordinates remediation activities", with "human approval and intervention where required" (FAQ). New Relic's SRE Agent entered preview on 24 February with a survey claim of "a 25% faster incident resolution time" (announcement). Splunk's AI SRE was slated GA for late June (Cisco Live post); Cisco's AI Canvas has the operator approve the plan "before any agent runs" (blog). Elastic completed its acquisition of Deductive AI on 24 August, an investigation agent with "a reinforcement learning harness for root cause analysis", terms undisclosed (release).

Product Status and date Price signal Remediation posture Source
Datadog Bits AI SRE GA 2 Dec 2025 6.5 credits ($6.50) per investigation; $500 per 500 credits a month Proposes; PR generation in private preview Datadog pricing, press release
PagerDuty SRE Agent GA core; virtual responder EA 4 AI Actions per request or nudge click Recommends; human clicks PagerDuty docs
Grafana Assistant Investigations GA 27 Jul 2026 In Grafana Cloud incl. free tier Investigates; Automations are scheduled prompts Grafana press release
Dynatrace Autonomous / Cloud SRE Agent Announced 27 Jul 2026, Aug availability Not disclosed Policy-driven approvals Dynatrace releases
New Relic SRE Agent Preview 24 Feb 2026 Not disclosed Recommends, can invoke workflows New Relic press release
Splunk AI SRE GA late June 2026 Not disclosed Plan and step-by-step guide Splunk blog
Cisco AI Canvas Controlled availability 2 Jun 2026 Not disclosed Operator approves plan before agents run Cisco blog
Rootly AI SRE Shipping 2026 Not disclosed "Every change requires explicit human sign-off" Rootly product page
Komodor Klaudia Platform launch 29 Sep 2026 Not disclosed Higher-risk actions can be held for approval ITBrief report

The startups and the money

The venture story is loud and the revenue story is quiet. Resolve AI raised a $35 million seed in October 2024 led by Greylock, then a Series A led by Lightspeed at a $1 billion headline valuation on 19 December 2025; TechCrunch's report put annual recurring revenue at roughly $4 million and the round size undisclosed. Traversal came out of stealth on 20 June 2025 with $48 million from Sequoia and Kleiner Perkins, claiming ">90% accuracy on hundreds of high-impact incidents", per its launch post. incident.io launched Investigations on 5 August 2026 (announcement); its earlier marketing promised MTTR cuts of "up to 80%". Cleric closed $5.5 million led by Vertex Ventures US in December 2025, $9.8 million total, with adopters "freeing 20 to 30% of engineering capacity" (release). Komodor, $90 million raised, relaunched as an agentic operations platform on 29 September where "higher-risk actions can be held for human approval" (report). Rootly's product page says "every change requires explicit human sign-off before execution"; RunWhen runs its MCP-exposed skills "with a human in the loop" (site). Parity, the reported Deductive price and the reported incident.io Series B I could not source from a primary page and have left out.

Notice the shape: everyone sells investigation, almost everyone who mentions remediation adds a sign-off clause, and the sector's largest valuation sits on a revenue base smaller than the seed round that preceded it. The buyers are still running pilots.

What is claimed versus what was measured

The vendor claims are consistent in form, hours to minutes, and inconsistent in evidence. The strongest numbers come from an operator: Google's SRE organisation reports that its Incident Hypothesis feature alone produced "a 10% reduction in Mean Time to Mitigate" and its Investigation Dashboards "roughly 44% reduction in Mean Time to Mitigate for supported incidents". Internal and unaudited, but the only MTTR deltas from a non-seller, and a calibration against "up to 80%".

The surveys measure sentiment. PagerDuty's 2026 State of AI-First Operations, 1,000 respondents, found 8 percent of organisations lose more than $1 million an hour in an incident, 34 percent at least $500,000 and 68 percent more than $300,000; 75 percent of AI adopters report improved resilience against 66 percent of the rest. An earlier PagerDuty study of 500 IT leaders put incidents up 43 percent in a year at $793,957 each. Uptime Institute's 2026 outage analysis is more sober: per-site outage frequency fell for the fifth year, 57 percent say their last major outage cost more than $100,000 and one in five more than $1 million, and "failures to follow established procedures remain the leading driver of human error-related outages". Its 2025 edition warned that software-based resiliency tools "can also introduce new risks and complexities".

The most interesting vendor evidence is Traversal's own benchmark of 14 August 2026: 25 real severity 1 to 3 incidents at a financial-services customer, blind pairwise Elo judging against frontier models given direct MCP or API access to the same stack. Traversal scored 1696 to 1441 for the best do-it-yourself setup, in 9.5 minutes against 19.7 and 33.5. A vendor document, but the point stands: the model is not the bottleneck, the harness and the data are.

The benchmarks, read without the press release

The academic benchmarks ask a narrower question, "can an agent find the root cause of an injected fault in a Kubernetes demo", and the answer has moved a lot without reaching anything you would page.

AI SRE benchmark scores, 2025 to 2026 (percent, best reported unless stated) strict (full root cause or verified fix) lenient or partial hard set 050%100% ITBench SRE resolved, Feb 2025 13.8 AIOpsLab mitigation, Jan 2025 54.5 (Flash agent) ITBench-AA, May 2026 (Opus 4.7) 47 ITBench-AA board, Oct 2026 (top) 56.2 OpenRCA 2.0 exact match, mean 20.7 OpenRCA 2.0 any correct service 76.0 SREGym end-to-end, May 2026 60.7 SREGym diagnosis only 72.6 ORCA-Bench medium, Jul 2026 25.3 ORCA-Bench hard 10.0 Incident-Arena, Sep 2026 (ceiling) below 64.3 Read across:strict metrics run 10 to 61 percent; the lenient ones (named a plausible service, found a symptom) run 72 to 76. ORCA-Bench's weakest model hallucinated an implausible root cause in 40 percent of reports. None of these is a production SLO.
Best reported scores from each benchmark's paper or leaderboard: ITBench (arXiv 2502.05352), AIOpsLab (arXiv 2501.06706), ITBench-AA (IBM and Artificial Analysis, May 2026 post and October 2026 leaderboard), OpenRCA 2.0 (arXiv 2606.27154), SREGym (arXiv 2605.07161), ORCA-Bench (arXiv 2607.28545), Incident-Arena (arXiv 2610.00648). Tasks, scoring and difficulty differ; the bars are not directly comparable.

IBM's original ITBench, February 2025, had the best agent resolving 13.8 percent of SRE scenarios and 0 percent of FinOps. Artificial Analysis's independent ITBench-AA, 27 May 2026, 59 Kubernetes tasks with 19 held out, scored Claude Opus 4.7 at 47 percent and GPT-5.5 at 46; scoring is zero if any ground-truth cause is missed. The leaderboard has since moved to 56.2 percent for GPT-5.6 Sol (Max), 43 points in twenty months. Microsoft's AIOpsLab found detection essentially solved while mitigation topped out at 54.5 percent.

The process-level work is more damning than the scores. OpenRCA 2.0, June 2026, annotated 500 incidents with causal paths and evaluated 11 frontier models: they named a correct root-cause service 76.0 percent of the time, grounded it in a verified path 61.5 percent of the time, and recovered the exact root-cause set 20.7 percent of the time. The authors name three patterns: premature commitment (stop at the first plausible cause), presence bias (silence read as health) and salience capture (fixate on the loudest downstream signal). SREGym, 90 live problems, has Claude Code with Sonnet 4.6 at 72.6 percent diagnosis and 60.7 percent end to end, dropping to 62.6 and 53.7 with realistic telemetry noise; on latent disk sector errors every agent scored zero on fault characterisation in every run, and when the diagnosis was wrong, mitigation still succeeded 35 to 59 percent of the time, which is restarting the pod and hoping. ORCA-Bench, 1,079 tasks over six days of real observability data, has the best model at 25.3 percent on the realistic setting and 10.0 on the hard one, the weakest model hallucinating "an implausible root cause in 40% of incident reports", and every model getting worse without source-code access; its authors include Traversal's founders. Incident-Arena, 30 September, puts frontier models below 64.3 percent on 20 repair tasks and tracks "unsafe regressions" as a failure class.

Benchmark What is scored Best result Source
ITBench (Feb 2025) SRE scenario resolved, 94 scenarios 13.8% arXiv 2502.05352
AIOpsLab (Jan 2025) Mitigation, 48 problems 54.5% (Flash); RCA 45.5% (ReAct) arXiv 2501.06706
ITBench-AA (May 2026) All root causes found, 59 tasks 47% Opus 4.7; 56.2% leader Oct 2026 IBM/AA post
OpenRCA 2.0 (Jun 2026) Exact root-cause set, 500 incidents 20.7% mean; 27.6% Opus 4.7 arXiv 2606.27154
SREGym (May 2026) Diagnosis / end to end, 90 problems 72.6% / 60.7% arXiv 2605.07161
ORCA-Bench (Jul 2026) RCA accuracy, 1,079 tasks 25.3% medium, 10.0% hard arXiv 2607.28545
Incident-Arena (Sep 2026) Verified repair, 20 tasks below 64.3% arXiv 2610.00648

Four outages where the automation was the incident

None of the big postmortems of the past twelve months involve an LLM. All involve automation that was correct, tested and in production.

AWS's account of 19 to 20 October 2025 in us-east-1 begins at 11:48 PM PDT with DynamoDB's automated DNS management: a Planner and redundant Enactors raced, an Enactor applied a stale plan after a newer one had cleaned it up, and the regional endpoint was left with an empty DNS record that "required manual operator intervention to correct". DNS was back at 2:25 AM, but EC2's workflow manager had entered "congestive collapse", and the Network Load Balancer's health-check automation kept removing capacity on alternating results until engineers disabled automatic failover at 9:36 AM; full recovery came at 1:50 PM. AWS "disabled the DNS Planner and Enactor automation worldwide" and is adding velocity controls so NLB cannot remove capacity faster than a human can notice.

Microsoft's post-incident review of Azure Front Door on 29 October 2025 is a safety system passing a bad change. A tenant configuration made at 15:35 UTC "completed propagation to a majority of edge sites by 15:39" because the protection system "continued to receive healthy signals"; the crash was global by 15:45, and because the fleet had reported healthy, the "last known good" snapshot was updated with the poison too. The fix was humans editing that snapshot by hand from 17:10 and blocking all customer configuration propagation at 17:30; stable at 00:05 UTC.

Cloudflare's 18 November 2025 failure, its worst since 2019, began with a ClickHouse permissions change that made a Bot Management query return duplicate rows; the feature file doubled past a hard-coded 200-feature limit and the Rust proxy panicked. Because the change rolled out gradually, the file regenerated every five minutes as sometimes good and sometimes bad, the fleet oscillated and responders spent hours suspecting a DDoS; core traffic recovered at 14:30. On 5 December a WAF change hit a nil-value bug in the older proxy and took out about 28 percent of HTTP traffic for 25 minutes; the configuration system, Cloudflare wrote, "does not perform gradual rollouts, but rather propagates changes within seconds to the entire fleet". Google Cloud's us-west1 incident on 20 August 2026, two hours and 22 minutes, followed fibre maintenance after which "automated rerouting mechanisms failed to properly redistribute traffic to alternate capacity"; humans drained traffic by hand.

Three things recur. The automation acted faster than any human could evaluate it. The health signals it trusted were wrong or late, exactly the data an agent would also read. And in every case the recovery step was a human turning the automation off. Add ORCA-Bench's 40 percent hallucinated causes and the risk of an agent that acts is not hypothetical. Change-freeze bypass is the one failure mode for which I found no public postmortem blaming an AI agent; the structural risk is visible in "within seconds to the entire fleet", and in Komodor's and Dynatrace's choice to make freeze-awareness a policy layer rather than trust the model to remember it. The agent that can open a PR can also merge one.

Investigators at full speed, a human on the button

The design I would accept, and the one most vendors have quietly converged on, is Google's L2 in its autonomy ladder: systems "can Monitor, Investigate, and even Actuate changes. However, a human must explicitly Approve any plan", with every proposed action risk-scored against deployments and error budgets. PagerDuty's nudges, Rootly's sign-off, Cisco's approve-before-run, Komodor's hold-for-approval and Datadog's review-and-merge PR are the same shape under different names. The investigator side should be unconstrained, parallel hypotheses, every data source, deploy correlation, past incidents, because that is where Google's 10 to 44 percent gains came from. The actuator side should be a short list of pre-approved, reversible, rate-limited actions with velocity controls of the kind AWS is now bolting onto NLB. incident.io's engineering write-up of 1 October adds the discipline most buyers skip: when they ask teams who built their own tooling how they measure accuracy, "the answer is usually that they don't"; theirs is backtesting every hypothesis against what responders actually did.

The cheapest investigator trick is also the one the agents lean on most: rank recent changes as suspects. I ran this in the scratchpad this week, stdlib only, on synthetic deploy events; it stops at a proposal.

# deploy_alert_correlation.py: rank recent changes as suspects for an alert, then STOP.
# Emits a proposal; nothing here calls kubectl, terraform or a cloud API.
# Python 3.14, stdlib only. Feed it your deploy log and the alert's onset time.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone

@dataclass
class Deploy:
    service: str
    change_id: str
    finished_at: datetime
    hops_from_alerting_service: int   # 0 = the service that paged, 1 = direct dependency, ...

def rank_suspects(alert_onset, deploys, window=timedelta(minutes=90)):
    """Score = how recently before onset the change landed, discounted by dependency distance.
    A change that landed after onset cannot have caused it and is dropped."""
    out = []
    for d in deploys:
        lead = alert_onset - d.finished_at
        if lead < timedelta(0) or lead > window:
            continue
        recency = 1.0 - lead / window                 # 1.0 = landed right before onset
        proximity = 1.0 / (1 + d.hops_from_alerting_service)
        out.append((round(recency * proximity, 3), d, lead))
    return sorted(out, key=lambda t: -t[0])

if __name__ == "__main__":
    onset = datetime(2026, 10, 5, 2, 14, tzinfo=timezone.utc)
    deploys = [
        Deploy("checkout-api", "chg-8812", onset - timedelta(minutes=11), 0),
        Deploy("pricing-svc",  "chg-8809", onset - timedelta(minutes=48), 1),
        Deploy("auth-gateway", "chg-8790", onset - timedelta(hours=5),    1),  # outside window
        Deploy("search-idx",   "chg-8815", onset + timedelta(minutes=3),  2),  # after onset
    ]
    for score, d, lead in rank_suspects(onset, deploys):
        print(f"{score:5.3f}  {d.service:<13} {d.change_id}  landed {int(lead.total_seconds()//60):>3} min before onset, {d.hops_from_alerting_service} hop(s)")
    print("PROPOSAL: roll back top suspect. Awaiting human approval; no action taken.")

Real output from the run, Python 3.14.6:

0.878  checkout-api  chg-8812  landed  11 min before onset, 0 hop(s)
0.233  pricing-svc   chg-8809  landed  48 min before onset, 1 hop(s)
PROPOSAL: roll back top suspect. Awaiting human approval; no action taken.

Thirty lines, and it already encodes two things OpenRCA says agents get wrong: it refuses to blame a change that landed after the symptom, and it discounts by dependency distance rather than by how loud the signal is. What the $48 million startups add is better data about what "hops" means in your system.

The ceiling is your telemetry, not the model

Every benchmark failure above has a data explanation before it has a model explanation. SREGym's agents scored zero on disk faults because the signal lives in kernel logs they were not reading; OpenRCA's presence bias is absent data read as health; ORCA-Bench's accuracy fell for every model when source code was withheld; incident.io admits scores are sometimes bad because "traces have fallen out of the retention period". Uptime's 2025 analysis attributed the rise in IT and network outages, 23 percent of impactful ones, to "change management and misconfigurations": deploy events nobody correlated with anything. An agent is a very fast reader of whatever you have written down. If your deploy events are not on the same clock as your alerts, if dependencies live in someone's head, if the runbook says "check the dashboard" without naming it, the agent inherits the gap and fills it with the most salient thing it can see. This is why agent observability and the boring discipline of structured deploy annotations matter more in the AI era, not less: the model is a commodity you rent for $6.50 an investigation, and the data it reads is the only part you own. The postmortems from my own fleet were, every time, about a signal I did not have rather than a decision I made badly.

What I take from the week. I would turn on an investigator today: for a solo operator, something that has read the runbook, ranked the deploys and tested three hypotheses before I find my glasses is worth the credits. I would not give it a kubectl that writes, and I would judge any vendor by how specific its approval gate is rather than by its MTTR headline. This quarter the work on my side is unglamorous: deploy events on the same timeline as alerts, a dependency map the agent can read instead of infer, retention long enough that the evidence outlives the postmortem, and a hard list of the five reversible actions I would let anything execute unasked. The agents will improve. The 2am question is still whether the machine should act on what it found, and the honest answer from every postmortem I read is that the humans were the ones who turned the automation off.


Related: Agent Observability in 2026, It's 2 AM. Do You Know What Your AI Agent Is Doing? and microVM Production Incidents: 4 Postmortems From My Fleet.

I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.

Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.


Related

Oct 5, 2026

AI SOC Agents in 2026: 98% Accuracy Claims, 23 to 34% on the Benchmark, and 0% of Teams Letting Them Act Alone

What the security-operations agents from Microsoft, Google, CrowdStrike, Palo Alto, SentinelOne, Torq and a billion-dollar startup cohort actually do in 2026 and what is measured: a median of 100 alerts a day and 28 percent never investigated, 75-minute mean investigations, Microsoft's $4-an-hour compute units and Google's token meter, CrowdStrike's 98 percent triage claim against Meta's 23 to 34 percent benchmark, Anthropic's and Google's reports of attackers running agents against defenders, and a Sigma rule and CI step that keep a human on the merge button.

14 min
Oct 5, 2026

CI/CD in 2026: Agents Opened 1M PRs in 5 Months, Bots Write 1 in 5 Reviews, and Most Agent PRs Get No Human Look

What the pipeline looks like when AI agents open the pull requests: GitHub's coding agent went GA on 25 September 2025 and opened over a million PRs in five months, Copilot code review passed 60 million reviews and one in five on GitHub, CodeRabbit has reviewed 13 million PRs on $88 million raised, and the first studies find most agent PRs get no human review attention. The gates that still hold: the assigner cannot approve, an extra approval for bot authors, path fences, signed commits, SLSA provenance, and a workflow YAML that gives agent PRs their own lane.

13 min
Oct 5, 2026

DevOps Still Matters in 2026: AI Cut Delivery Stability 7.2%, Then Doubled Merged PRs and Added 91% to Review Time

Why DevOps is the big thing of the agent era: DORA 2024 found a 25% rise in AI adoption cost 1.5% throughput and 7.2% stability, DORA 2025 saw throughput turn positive while instability stayed up across nearly 5,000 respondents, and DORA's 2026 ROI model budgets a 15% three-month dip and a change failure rate rising from 5% to 6%; GitHub merged 518.7M PRs (+29%) and over 1M agent PRs in five months, Faros telemetry on 10,000 developers shows 98% more PRs and 91% longer reviews, METR found experienced developers 19% slower, and the Replit postmortem's fixes are 2015 DevOps controls.

13 min