AI SOC Agents in 2026: 98% Accuracy Claims, 23 to 34% on the Benchmark, and 0% of Teams Letting Them Act Alone
The security alerts for my platform land in a queue that I read, because there is no one else to read them. That is the state of security operations for most small teams, and it is why I have watched the "AI SOC agent" market more closely than I usually watch a Gartner category. The pitch: an agent reads every alert, investigates the way a tier-one analyst would, closes the noise and hands a human the handful that matter. In 2026 the pitch is backed by shipped products, real money and a few real numbers, and by many numbers that are not measurements. This post sorts them: what the agents actually do and how they are priced and governed, what the surveys say about the queue, what attackers are doing with the same models, and what the only independent benchmarks say about the models underneath. I run PandaStack, a Firecracker microVM cloud for agent workloads, so I have an interest in agents being given safe places to run; I have no stake in any security vendor named here.
The queue the agents are being hired for
The best recent measurement of the problem is the second State of AI in the SOC survey, run by Prophet Security with ViB and published on 3 August 2026 from 250 security leaders and practitioners. It is vendor-sponsored, so weigh the adoption figures accordingly, but the workload numbers are the useful part. The median organisation takes in about 100 alerts a day and the largest about 1,000; 74 percent receive 50 or more and 27 percent receive 500 or more. Mean time to investigate one alert is 75 minutes, median 45. The average organisation leaves roughly 28 percent of its alerts uninvestigated, and 60 percent say an alert they ignored later turned out to matter. The 2025 edition (282 respondents) had 40 percent uninvestigated and 960 alerts a day; two vendor surveys are not a trend.
SANS's numbers point the same way. Its 2025 Detection and Response survey, summarised by Stamus Networks, has 73 percent of organisations naming false positives as their top detection challenge, with "very frequent" false positives up from 13 to 20 percent in a year. SANS's 2025 AI survey found only half of organisations using AI for any security task, 33 percent for investigation, and 66 percent complaining that the AI systems themselves generate excessive false positives; its 2025 SOC survey, as Swimlane reports it, has 69 percent of SOCs still reporting metrics by hand. Burnout figures of 70-plus percent circulate from several vendors; I could not reach a primary source for any and have left them out.
The number I keep coming back to is the last bar. Among the 40 percent running AI in production, 57 percent require a human to review the verdict before an alert is closed, 44 percent let the agent recommend only, and nobody grants full autonomy. Only 30 percent say the agent's verdicts match an experienced analyst at least 90 percent of the time; 22 percent put agreement at 50 to 69 percent, a coin flip with a confidence score. And 72 percent of users tried building their own tooling first, of whom 46 percent scrapped it.
What the platform vendors shipped
Microsoft's agent is the most documented, which is to its credit. The Phishing Triage Agent was announced on 24 March 2025 with preview in April, and the Defender documentation, updated August 2026, now calls it the Security Alert Triage Agent: triage of user-reported email is generally available, and a preview extends it to a list of Defender for Cloud container alerts and three identity alerts. It classifies an alert True or False Positive, resolves the false positives, leaves the true positives open for a human, and accepts natural-language feedback that an administrator must approve before it becomes a "lesson" in the agent's memory. It does not isolate hosts or disable accounts. Pricing is in Security Compute Units: $4 an hour per provisioned SCU, $6 for overage, and the Microsoft 365 E5 and E7 inclusion announced at Ignite gives 400 SCUs a month per 1,000 licences, capped at 10,000, rolling out from 18 November 2025. The usage dashboard reports cost per email processed, which is the right unit; Microsoft's own FAQ says a manually triaged phish takes up to 30 minutes.
Google's Alert Triage and Investigation Agent entered public preview on 12 November 2025 for SecOps Enterprise and Enterprise Plus. It runs an investigation plan drawn from Mandiant practice and returns a verdict with a confidence score; it recommends, it does not remediate. By June 2026 Google said it had investigated more than 5 million alerts and cut a typical 30-minute analysis to 60 seconds, with no false-positive rate published; I could not reach the original post and rely on that summary. The release notes add a Threat Hunt Agent in preview on 3 August 2026, a Detection Engineering Agent on 18 August that drafts YARA-L rules over MCP tools, and, from 1 July 2026, "Security Tokens" as the meter for agentic consumption. Both hyperscalers now sell SOC agents by the compute unit, which is the first line in your FinOps model.
CrowdStrike has the boldest accuracy number. Its Charlotte AI product page claims over 98 percent decision accuracy for Agentic Detection Triage, benchmarked against the verdicts of its own Falcon Complete team, and states that Charlotte does not execute response actions by default: each Agentic SOAR workflow is set to autonomous or approval-required, with credit caps and execution traces. Seven prebuilt agents and the AgentWorks builder arrived on 16 September 2025, and at Fal.Con on 2 September 2026 the investigation layer was rebuilt to run agents across endpoint, identity, SaaS, cloud and network in parallel, alongside an identity provider issuing short-lived, scoped credentials to agents (see the non-human identity post).
Palo Alto Networks turned XSOAR into Cortex AgentiX on 28 October 2025: six prebuilt agents, more than 1,000 integrations, claims of up to 98 percent lower mean time to respond, and human approval for impactful actions. In February 2026 XSIAM gained Case Investigation, Cloud Posture and Automation Engineer agents, and the May 2026 release put "Agentic Response" into XSIAM 3.5 with approvals for sensitive actions. SentinelOne's Purple AI Athena, unveiled at RSAC on 29 April 2025, ships Auto-Triage as generally available with investigation, rule creation and remediation in preview, works on third-party SIEM data, and its release post publishes no accuracy figures.
| Vendor and product | What it does today | Autonomy and pricing | Published accuracy | Source |
|---|---|---|---|---|
| Microsoft Security Alert Triage Agent | Classifies user-reported phish (GA), cloud and identity alerts (preview); resolves FPs | Classify only; $4/SCU-hour, $6 overage; E5/E7 inclusion 400 SCUs per 1,000 seats | None published; "measurable improvements in controlled evaluations" | Docs |
| Google SecOps Triage and Investigation Agent | TP/FP verdict with confidence, investigation plan; Threat Hunt and Detection Engineering agents in preview | Recommends only; Security Tokens meter from 1 July 2026 | 5M+ alerts, 30 min to 60 s; no FP rate | Preview post |
| CrowdStrike Charlotte AI | Detection triage verdict and confidence; parallel multi-domain investigation; Agentic SOAR | Per-workflow autonomy level; no response by default; credit caps | >98% vs Falcon Complete analysts (vendor) | Product page |
| Palo Alto Cortex AgentiX / XSIAM | Six-plus prebuilt agents, Agentic Response in XSIAM 3.5 | Human approval for impactful actions | Up to 98% MTTR reduction (vendor); Tyson 50% | Press release |
| SentinelOne Purple AI Athena | Auto-Triage GA; investigation, rules, remediation preview; third-party SIEMs | Recommendations reviewed by analysts | None published | Blog |
| Torq HyperSOC / Socrates | Omni-agent investigation and response over workflows | Guardrails per workflow | Carvana: 100% of Tier-1 alerts; a retail CISO: over 50% closed autonomously | Series D |
| Intezer AI SOC | Forensic triage of 100% of alerts across EDR, identity, email, cloud | Escalates <2% to humans | 98% verdict accuracy (vendor) | Product |
The money, and the Gartner curve
The independent cohort raised at a pace that says where investors think tier-one work is going. Torq closed a $140 million Series D on 12 January 2026 at a $1.2 billion valuation, $332 million raised in total. 7AI, founded by Cybereason alumni, raised a $130 million Series A led by Index Ventures on 4 December 2025 and reported 7 million investigations in production by May 2026. Exaforce announced a $125 million Series B on 12 May 2026, $200 million in total, led by Mayfield, Khosla, HarbourVest and Peak XV among others. Prophet Security's roughly $30 million Series A led by Accel in July 2025 I have only from Tracxn, so treat it as reported. Intezer said on 16 September 2026 that first-half revenue tripled. I could not reach primary pages for Dropzone AI's funding and have left its numbers out; Gartner lists it among its sample vendors.
Gartner's position is cooler than the funding. In the Hype Cycle for Security Operations dated 5 June 2026, as Dropzone summarises it, AI SOC agents sit at the Peak of Inflated Expectations with 1 to 5 percent penetration and "embryonic" maturity. An earlier note, reported by BleepingComputer in March 2026, predicts that 70 percent of large SOCs will pilot agents for tier-one and tier-two work by 2028 and only 15 percent will show measurable improvement without a structured evaluation.
What the benchmarks say about the models underneath
Every accuracy figure above is a vendor measuring its own product against its own analysts on its own alerts. The only public, reproducible measurements are of the models, not the products. CyberSOCEval, from Meta and CrowdStrike, posted on 24 September 2025 as part of CyberSecEval 4, asks 609 malware-analysis questions built from Falcon Sandbox detonations and 588 threat-intelligence questions drawn from 45 real reports, scored strictly: a question counts only if the model picks all the correct options and none of the wrong ones. Frontier models from OpenAI, Google, Anthropic, Meta and DeepSeek score 23 to 34 percent on malware analysis against a 0.6 percent random baseline, and 43 to 53 percent on threat-intelligence reasoning against 1.7 percent. Reasoning models gain little over their base models, which the authors read as evidence that nobody has trained these models to reason about security analysis. The paper predates the current model generation, and says so.
Simbian's June 2026 cyber-defence benchmark is vendor-run but tests the actual job: 27 models, 1,132 runs over 26 multi-stage attack campaigns, each investigated by SQL over 75,000 to 135,000 log records with a budget of 50 queries, scored by how much of the attack narrative the model reconstructs across the 13 ATT&CK tactics. The best model surfaced 45.1 percent of the coverable steps at $3.76 a run; none of the 27 cleared a 50 percent bar on every tactic, Simbian's threshold for unsupervised operation. The gap between "98 percent of the verdicts our analysts would have given on phishing reports" and "45 percent of an attack chain reconstructed from raw logs" is not a contradiction. It is the difference between classifying a bounded, high-volume alert class and investigating an incident, and the honest products scope themselves to the former.
The other side is running agents too
The SOC cannot simply wait for the models to improve, because the attackers are not waiting. Anthropic's August 2025 report described a single operator using Claude Code to extort at least 17 organisations with demands sometimes above $500,000. The November 2025 report described a state-backed campaign against roughly 30 targets in which the model did 80 to 90 percent of the work with human intervention at four to six decision points per intrusion. The June 2026 analysis mapped 832 banned accounts onto ATT&CK: 67.3 percent were developing malware. The September 2026 report, covering December to August, has the case every detection engineer should read: a Russian state-nexus actor whose agents watched security products for detections of its own implants and then autonomously rebuilt the malware until it stopped being detected, against more than 20 targets. The same report has a criminal crew going from one stolen developer token to full cloud administrator access in about three hours, and a Chinese operation producing more than a dozen possible zero-day findings a month from automated binary analysis, the hostile mirror of the Glasswing numbers in my vulnerability-discovery post.
Google's Threat Intelligence Group has the matching record. Its November 2025 tracker documented PROMPTSTEAL, deployed by APT28 against Ukraine, generating its Windows commands at runtime from an open-weight coder model; PROMPTFLUX, an experimental dropper asking Gemini for fresh evasion code every hour; and FRUITSHELL, a reverse shell carrying prompts meant to confuse LLM-based scanners. The September 2026 tracker, "From Prompting to Autonomy", describes an actor compromising cloud resources and then planning, building and executing an agent-driven credential-harvesting campaign in under six hours, and a credential stealer that embeds deliberately policy-violating text in its loaders so that an LLM-based scanner refuses to analyse the file and moves on. The same malware hides payloads in the dotfile directories of coding assistants and editors so they execute when a developer opens the repository. That is a detection-engineering problem with a clear shape, so I wrote the rule.
Detection as code, with the human on the merge
If agents are going to write detections, and Google's and CrowdStrike's now do, the pipeline that reviews, tests and compiles them matters more than the model that drafted them. Sigma remains the vendor-neutral format, with more than 3,000 rules under the Detection Rule License 1.1 and roughly monthly releases (r2026-07-01 shipped on 9 July 2026). I wrote a rule for the dotfile-directory persistence above and ran it through the toolchain in a scratch venv today; pip resolved sigma-cli 3.1.0, pySigma 1.5.1 and the Splunk backend 2.1.0. The first sigma check rejected the rule for a malformed ATT&CK tag, exactly the mistake a model drafting rules at volume will make and exactly what the validator exists to catch. After the fix it passed with zero issues.
# rules/proc_creation_lnx_ai_tool_workspace_dirs.yml
# Shell or interpreter launched with a command line that points into the hidden
# config directories of coding assistants and editors. GTIG's September 2026
# tracker describes malware parking payloads there so they run when the IDE or
# AI extension opens the workspace.
title: Shell Spawned From AI Tool Workspace Config Directory
id: 3f1c2e7a-9b4d-4c1e-8f2a-6d5b7c8e9a01
status: experimental
references:
- https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai
author: Ajay Kumar
date: 2026-10-05
tags: [attack.persistence, attack.t1546]
logsource: {category: process_creation, product: linux}
detection:
selection_shell:
Image|endswith: ['/bash', '/sh', '/zsh', '/python3', '/node']
selection_path:
CommandLine|contains: ['/.claude/', '/.cursor/', '/.vscode/']
filter_known_good: # the tools themselves legitimately run tasks from here
ParentImage|endswith: ['/code', '/cursor', '/claude']
condition: selection_shell and selection_path and not filter_known_good
falsepositives: [Developers running project scripts kept under .vscode by hand]
level: medium
# .github/workflows/detections.yml
# Every rule change, human- or agent-authored, is validated and compiled in CI.
# An agent may open the pull request; branch protection means it cannot merge it.
name: detection-as-code
on: {pull_request: {paths: ['rules/**']}}
jobs:
validate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: {python-version: '3.12'}
- run: pip install sigma-cli==3.1.0 pysigma-backend-splunk==2.1.0
- run: sigma check rules/ # schema, conditions, ATT&CK tags
- run: sigma convert -t splunk -p splunk_cim rules/ > build/rules.spl
# replay build/rules.spl against tests/positive and tests/negative log fixtures here;
# a rule that fires on the negative set fails the build, no matter who wrote it
The convert step emitted a Splunk CIM search over Processes.process_path, Processes.process and Processes.parent_process_path with the same three clauses. The false-positive filter is the part a model cannot write for you, because it depends on what your developers actually do, and it is the part that decides whether the rule adds to the 73 percent problem or subtracts from it.
Where to let the agent act
Every vendor above defaults to classify-and-recommend; the ones that offer containment gate it behind a per-workflow switch; the surveyed teams let the agent close false positives and recommend the rest, and none let it act alone. That is not timidity. An agent that closes a false positive wrongly costs one missed alert, which is the 28 percent you were already missing. An agent that isolates a production host or disables a service account wrongly costs an outage and an awkward retrospective, and the attacker reports show adversaries now crafting inputs for LLM-based tools, so the agent's own decision is an attack surface. Split by blast radius: autonomous closure for high-volume, well-bounded alert classes where the agent's agreement with your analysts has been measured on your alerts, not the vendor's; recommended actions with one-click approval for the rest; and no path from the model to a destructive API without a human or a deterministic policy in between, the same shape I argued for from inside the sandbox in the 2 AM post.
What I take from all of it. The agents are real, they are priced by the compute hour or the token, and for the one alert class with a published method, user-reported phishing, closing false positives at machine speed is clearly worth four dollars an hour. Beyond that class the only numbers are vendors' own, the public benchmarks put the models at a quarter to a half of a passing score on analysis tasks, and the attackers have already adapted to being read by a model. This quarter I would do three things in order: put every detection I own into a repository with the validate-and-compile pipeline above so that agent-drafted rules have somewhere safe to land; turn on a triage agent for the one alert type I can measure against my own verdicts and log its agreement rate for ninety days before I believe a percentage; and keep every containment action behind an approval I have to click, while the surveys still report that nobody has taken their hand off it.
Related: AI Found the Bugs: 23,000 Findings in a Month, Who Is This Agent? Non-Human Identity in 2026 and AIOps in 2026: Incident Response Agents and MTTR.
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
AIOps in 2026: 47% on ITBench, 10% on Hard On-Call, $6.50 an Investigation, and Four Outages Where the Automation Was the Incident
The incident-response agents of 2026 against the evidence: Datadog Bits AI SRE GA at about 6.5 credits ($6.50) per investigation, PagerDuty's approval-gated SRE Agent, Grafana's six GA agent tools, Splunk's AI SRE, Elastic buying Deductive, Resolve at a $1B headline valuation on $4M ARR, Traversal's $48M; benchmarks from 13.8% on ITBench in 2025 to 47% in May 2026, 20.7% exact root cause on OpenRCA 2.0, 10% on hard ORCA-Bench and 40% hallucinated causes; what the AWS, Azure, Cloudflare and Google postmortems say about automation; and why a human should hold the button.
14 minOct 5, 2026CI/CD in 2026: Agents Opened 1M PRs in 5 Months, Bots Write 1 in 5 Reviews, and Most Agent PRs Get No Human Look
What the pipeline looks like when AI agents open the pull requests: GitHub's coding agent went GA on 25 September 2025 and opened over a million PRs in five months, Copilot code review passed 60 million reviews and one in five on GitHub, CodeRabbit has reviewed 13 million PRs on $88 million raised, and the first studies find most agent PRs get no human review attention. The gates that still hold: the assigner cannot approve, an extra approval for bot authors, path fences, signed commits, SLSA provenance, and a workflow YAML that gives agent PRs their own lane.
13 minOct 5, 2026CRA Reporting Went Live on 11 September: 24 Hours to ENISA, 3.08 Billion Rekor Entries, 17% of PyPI Attested, 92% SBOM False Positives
Supply-chain compliance in 2026: what the Cyber Resilience Act's Article 14 duty requires since 11 September 2026 (24-hour early warning, 72-hour notification, ENISA's Single Reporting Platform), what waits until 11 December 2027, how the open-source steward role works, where SBOMs stand (CISA's 2026 minimum elements, CycloneDX 1.7, a 92 percent false-positive rate in a 2,414-repo study), how far provenance has got (SLSA 1.2, 20 percent of PyPI uploads via trusted publishing, Rekor at 3.08 billion entries), the US retreat from mandates, and the pipeline I would run.
13 min