DevOps Still Matters in 2026: AI Cut Delivery Stability 7.2%, Then Doubled Merged PRs and Added 91% to Review Time

Oct 5, 2026 · 13 min · Ajay Kumar

A good share of the pull requests I merge into PandaStack now start life in an agent, and I am the only reviewer. (PandaStack is my company, a Firecracker microVM cloud for agents and code execution.) That is a strange position for someone who spent fourteen years telling teams that the pipeline, not the typing, was the constraint, because this year the typing became nearly free and the constraint became impossible to ignore: every agent PR still needs tests, a scan, a human read, a deploy window, a rollback path and someone awake if it goes wrong. None of those got faster. The industry's measurements say the same with more rigour than my commit log, so this post goes through what DORA found across three reports, what GitHub, METR, Stack Overflow and Faros measured, which 2025 incidents were about AI, and what I would gate this quarter.

The volume claims, and what they count

Sundar Pichai told Alphabet's April 2025 earnings call that AI generated more than 30 percent of Google's new code; Satya Nadella, at Meta's LlamaCon on 29 April 2025, put Microsoft at 20 to 30 percent "written by software", with better results in Python than C++. Dario Amodei told the Council on Foreign Relations on 10 March 2025 that AI would be writing 90 percent of code within three to six months and "essentially all of the code" within twelve. Anthropic's own survey of 132 of its engineers, published 2 December 2025, is more careful: Claude is used in 59 percent of their work, up from 28 percent a year earlier, the self-estimated productivity boost is 50 percent, and engineers say they can fully delegate between zero and 20 percent of it.

The volume is real at the repository level. GitHub's Octoverse 2025, published 28 October 2025, counts 518.7 million merged pull requests in the year, up 29 percent, 43.2 million a month, 986 million commits, up 25.1 percent, and over one million pull requests opened by the Copilot coding agent between May and September 2025, concentrated in older, well-starred repositories.

None of these is a delivery metric. A share of characters typed by a model, or a count of PRs opened, measures output at the keyboard and says nothing about whether the change shipped, broke anything, or sat for a week waiting for a reviewer. Those are the questions DORA has asked for a decade.

Three DORA reports, one direction

The 2024 Accelerate State of DevOps report, announced on 23 October 2024, set the tone. For every 25 percent increase in AI adoption it estimated a 7.5 percent improvement in documentation quality, 3.4 percent in code quality and 3.1 percent in code review speed, and alongside those a 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability. More than 75 percent of respondents relied on AI for at least one daily task; 39 percent reported little or no trust in its output. The same report found a parallel paradox for platform engineering, per OpsLevel's summary: an 8 percent gain in individual productivity and 10 percent in team performance, but an 8 percent drop in throughput and a 14 percent drop in change stability where teams used the platform exclusively. In 2024 both of the industry's favourite tools made individuals faster and the system less stable.

The 2025 report, renamed the State of AI-assisted Software Development, with findings from nearly 5,000 respondents published 23 September 2025, moved in one respect. Ninety percent now used AI, up 14 points, for a median two hours a day; more than 80 percent said it improved their productivity, yet only 24 percent trusted its output "a great deal" or "a lot" against 30 percent "a little" or "not at all". At the organisational level AI was, for the first time, associated with higher throughput, but still with higher instability and no reduction in friction or burnout, as summaries of the report put it. DORA's response was the AI Capabilities Model, seven conditions under which AI helps: a clear and communicated AI stance, healthy data ecosystems, AI-accessible internal data, strong version control practices, working in small batches, a user-centric focus and quality internal platforms. Four of the seven are DevOps practices with a new label, and platforms are described as the governance layer that turns individual gains into organisational ones.

For 2026, what DORA has published is the ROI of AI-assisted Software Development report, dated 22 April 2026 on dora.dev and covered by InfoQ on 11 May. It introduces a J-curve, a 15 percent productivity dip for about three months before gains arrive, and names the cause the "verification tax": the effort to check that generated code is reliable, secure and consistent with the architecture. Its worked model for a 500-person engineering organisation assumes a 12.5 percent net time gain after the dip, deployments rising from 50 to 56 a year, and a change failure rate rising from 5 to 6 percent, which at $100,000 an hour of downtime costs $344,000 a year; it nets $11.6 million of value against $8.4 million of investment, a 39 percent first-year return. The sentence it repeats is "AI will not fix broken engineering systems". As of 5 October 2026 dora.dev's publications page lists no 2026 annual report; the 2025 edition arrived on 23 September, so this year's is late or renamed, and nothing from it is quoted here.

Report Date Sample AI and throughput AI and stability Other Source
2024 Accelerate State of DevOps 23 Oct 2024 not stated in announcement -1.5% per 25% more adoption -7.2% per 25% more adoption +7.5% docs, +3.4% code quality, +3.1% review speed; 39% little or no trust Google Cloud
2025 State of AI-assisted Software Development 23 Sep 2025 ~5,000 positive, first time still negative; no less friction or burnout 90% adoption, 2 h/day median, 24% high trust vs 30% low; seven capabilities Google
2026 ROI of AI-assisted Software Development 22 Apr 2026 model, not survey 15% dip for 3 months, then +12.5% time; 50 to 56 deploys/yr change failure rate 5% to 6% 39% first-year ROI, 8-month payback, "verification tax" dora.dev, InfoQ
DORA 2024: what AI and platforms did to the numbers (percent change) improves worsens AI adoption, per 25% increase documentation quality +7.5 code quality +3.4 code review speed +3.1 delivery throughput -1.5 delivery stability -7.2 Internal platform, exclusive use individual productivity +8 throughput -8 change stability -14 -14-70+7 2025, ~5,000 respondents: AI now linked to higher throughput, still higher instability, no less burnout 2026 ROI model: 15% dip for three months, change failure rate 5% to 6%; no 2026 annual report as of 5 Oct
Effect sizes from the 2024 DORA report as stated in Google Cloud's announcement and OpsLevel's summary; the footer summarises the 2025 report and the 2026 ROI model. Scale: 14 units per percentage point.

The individual speed-up is smaller than it feels

DORA's numbers are survey-based, so it matters that the one randomised trial points the same way. METR's study, published 10 July 2025, gave 16 experienced open-source maintainers 246 real issues in their own repositories, averaging over 22,000 stars and a million lines, randomised each task to AI-allowed or AI-forbidden with Cursor Pro and Claude 3.5 and 3.7 Sonnet. The AI-allowed tasks took 19 percent longer. The developers had forecast a 24 percent speed-up and, after the fact, still believed they had been 20 percent faster. METR's February 2026 update widened the panel to 57 developers, 143 repositories and over 800 tasks using late-2025 agentic tools such as Claude Code and Codex: the original ten developers measured 18 percent slower with a confidence interval from 38 percent slower to 9 percent faster, and the 47 new recruits 4 percent slower with an interval from 15 percent slower to 9 percent faster. METR calls the new signal unreliable, because the most AI-optimistic developers stopped volunteering tasks, and is redesigning the experiment. For experienced engineers on code they know, the measured effect sits between a modest slowdown and a modest speed-up, and the perceived effect is reliably more flattering.

Developers' own trust tracks the measurements, not the perception. Stack Overflow's 2025 survey of 33,662 developers found 84 percent using or planning to use AI tools, while only 3.1 percent highly trusted the output and 45.7 percent distrusted it; 66 percent named "almost right, but not quite" solutions as their biggest frustration and 45.2 percent said debugging AI code takes longer than writing it. Only 30.9 percent used agents at all. The 2026 survey had not been posted on the survey site when I checked, so the numbers circulating for it are not here.

Where the queue forms

If individuals are not much faster but teams merge far more code, something is absorbing the difference. Faros AI's telemetry study of 28 July 2025, covering over 10,000 developers in 1,255 teams, shows where. Teams with high AI adoption completed 21 percent more tasks and merged 98 percent more pull requests, but their PRs were 154 percent larger, code review took 91 percent longer, and bugs per developer rose 9 percent. At the organisational level the study found no correlation between AI adoption and cycle time. The work did not get through faster; it piled up at review, in bigger pieces, with more defects.

This is the Theory of Constraints happening in public: speeding up a non-bottleneck does not raise throughput, it raises inventory in front of the bottleneck. In 2024 the inventory showed up as a 7.2 percent stability cost. In 2025 teams pushed harder on review and deploy, throughput turned positive, and instability stayed. In 2026 DORA gave the inventory a name, the verification tax, and a shape, the J-curve. Octoverse offers one hopeful datum, 72.6 percent of Copilot code review users said it improved their effectiveness, so part of the review bottleneck may be automated. But automated review of automated code is a pipeline, and a pipeline is a DevOps artefact: built, tested, measured by change failure rate, and owned by someone.

The incidents that were about AI, and the ones that were not

The incident everyone cites is Replit. Between 12 and 20 July 2025 Jason Lemkin, the SaaStr founder, built an app with Replit's agent, declared a code freeze, and told the agent, in his words to The Register, "eleven times in ALL CAPS" not to change anything; the agent then deleted the production database, created a 4,000-record database of fictional people, produced false test results, and told him rollback was impossible when it was not. Amjad Masad called it "unacceptable and should never be possible" and listed the fixes: automatic separation of development and production databases, additional staging environments, forced lookup of Replit's own documentation, and a one-click restore of project state. Environment separation, staging, tested backups and a change freeze enforced by policy rather than by prompt are the DevOps controls of 2015. The agent did not create a new class of failure; it removed the human who used to supply those controls by habit.

The two largest outages of late 2025 were not about AI, which is the more important point. Cloudflare's 18 November postmortem traces three hours of global failure, 11:28 to 14:30 UTC, to a permissions change at 11:05 that made a ClickHouse query return duplicate columns, doubling a Bot Management feature file past its 200-feature limit as it propagated. AWS's summary of the 19 to 20 October event traces it to a race between two DNS automation workers for DynamoDB's us-east-1 endpoint that left an empty record from 11:48 PM to 2:40 AM Pacific, with dependent services recovering for hours afterwards. Neither document mentions AI tooling. Both are failures of configuration propagation, change control and automation safety, the paths agent-generated volume now flows through. If a human's permissions change can take down much of the web for three hours, the discipline around changes is the thing to invest in, not who typed them. I wrote about what agents do unsupervised in It's 2 AM. Do You Know What Your AI Agent Is Doing? and why isolation is not authorisation in Isolation Is Not an Abuse Control. A widely repeated story that an AWS agent caused an internal outage in December 2025 I could not verify from any primary source, so it is not here.

Gating the agent's pull request

GitHub already encodes two of the right rules for its own agent: its documentation says that "by default, GitHub Actions workflows will not run automatically when Copilot pushes changes to a pull request" until someone with write access approves them, and that if you asked Copilot for the change, "your approval of a Copilot pull request won't count"; another reviewer must approve. The workflow below makes the same policy explicit for any agent and adds two things DORA's capabilities model asks for that GitHub does not enforce: a batch-size cap and a security scan the agent cannot talk its way past. I validated that it parses and has the job graph I intended, not that it runs against a live repository; treat the test and dependency steps as a template.

# .github/workflows/agent-pr-gate.yml
# Gates pull requests opened by coding agents: tests, CodeQL, a batch-size cap,
# and an approval from a human who is not the PR author or the person who prompted the agent.
# GitHub already refuses to run Actions on Copilot-agent pushes until someone clicks
# "Approve and run workflows", and the requester's approval does not count toward
# required reviews; this workflow makes the same rules explicit for any agent.
name: agent-pr-gate
on:
  pull_request:
    types: [opened, synchronize, reopened, labeled]
  pull_request_review:
    types: [submitted]

permissions:
  contents: read
  pull-requests: read
  security-events: write     # CodeQL upload

env:
  MAX_CHANGED_LINES: "400"   # DORA: work in small batches; large agent PRs sit in review longest

jobs:
  classify:
    runs-on: ubuntu-24.04
    outputs:
      agent: ${{ steps.who.outputs.agent }}
    steps:
      - id: who
        # Treat as agent-authored if the actor is a known bot, the branch uses an agent prefix,
        # or a human labelled it. Extend the list for your own agents.
        run: |
          case "${{ github.event.pull_request.user.login }}" in
            copilot-swe-agent[bot]|devin-ai-integration[bot]|claude[bot]) agent=true ;;
            *) agent=false ;;
          esac
          case "${{ github.head_ref }}" in copilot/*|agent/*|claude/*) agent=true ;; esac
          if printf '%s\n' '${{ toJson(github.event.pull_request.labels.*.name) }}' | grep -q '"agent-authored"'; then agent=true; fi
          echo "agent=$agent" >> "$GITHUB_OUTPUT"

  batch-size:
    needs: classify
    if: needs.classify.outputs.agent == 'true'
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@v4
        with: { fetch-depth: 0 }
      - run: |
          lines=$(git diff --numstat origin/${{ github.base_ref }}...HEAD | awk '{a+=$1+$2} END {print a+0}')
          echo "changed lines: $lines"
          test "$lines" -le "$MAX_CHANGED_LINES" || { echo "::error::agent PR exceeds $MAX_CHANGED_LINES changed lines; split it"; exit 1; }

  tests:
    needs: classify
    if: needs.classify.outputs.agent == 'true'
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -r requirements.txt -r requirements-dev.txt
      - run: pytest -q --maxfail=1       # the agent's own "tests pass" claim is not evidence

  sast:
    needs: classify
    if: needs.classify.outputs.agent == 'true'
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@v4
      - uses: github/codeql-action/init@v3
        with: { languages: python, queries: security-extended }
      - uses: github/codeql-action/analyze@v3
        with: { category: agent-pr }

  human-approval:
    needs: [classify, batch-size, tests, sast]
    if: needs.classify.outputs.agent == 'true'
    runs-on: ubuntu-24.04
    steps:
      - env:
          GH_TOKEN: ${{ github.token }}
          PR: ${{ github.event.pull_request.number }}
          AUTHOR: ${{ github.event.pull_request.user.login }}
        # Count APPROVED reviews from accounts that are not bots and not the PR author.
        # Pair this with a branch ruleset requiring 1 approval and dismissing stale reviews.
        run: |
          n=$(gh api "repos/$GITHUB_REPOSITORY/pulls/$PR/reviews" --paginate \
              --jq "[.[] | select(.state==\"APPROVED\" and .user.type==\"User\" and .user.login!=\"$AUTHOR\")] | length")
          echo "human approvals: $n"
          test "$n" -ge 1 || { echo "::error::agent-authored PR needs an approval from a human who did not author it"; exit 1; }

The important design choice is the classify job: once agent PRs are a labelled class, you can compute DORA's four keys for that class separately, which is the measurement the 2026 ROI model assumes you have and almost nobody does. If agent changes show the 5-to-6-percent failure-rate drift the model predicts, you have found your verification tax; if they do not, you have evidence to loosen the cap.

The platform jobs did not disappear

Dynatrace's State of SRE and Platform Engineering 2026, from 919 senior IT leaders surveyed between October 2025 and January 2026, finds 89 percent of organisations practising platform engineering running an internal developer platform, 73 percent of SRE and platform teams sharing responsibilities, and 89 percent using SLOs for at least some systems. The roles' AI content shifted rather than shrank: 67 percent of SREs name AI model monitoring their top use case, 50 percent use AI for automated incident response, and 55 percent of platform teams say their priority is enabling developers with AI coding tools, while the same leaders say AI is delivering less than expected on cost and time to recover. Indeed's Hiring Lab reported on 17 September 2026 that advertised pay in the most AI-exposed occupations has risen about 46 percent since 2021 against 25 percent in the least exposed, and that in those occupations the entry-level share of postings fell from 29 percent to 10 percent while the senior share rose from 22 to 47 percent. The market pays more for fewer, more senior people who can own a system, which describes a platform engineer. Gartner's much-quoted forecast of 80 percent of large engineering organisations having platform teams by 2026 I could not retrieve from Gartner's own site, and I found no primary count of "AI SRE" postings, so neither is in my ledger. The observability side, the spans and evals that make an agent's behaviour reviewable at all, I covered in Agent Observability in 2026.

What I take from all of it is that the thesis "AI makes DevOps obsolete" has the causality backwards. Generation got cheap, so the cost of a change moved to the steps that were always expensive: integrating, verifying, delivering, operating, and knowing when to roll back. DORA measured the bill in 2024, watched teams start paying it in 2025, and in 2026 told them to budget for it. This quarter, on my own platform, that means four things: compute the four keys separately for agent PRs, cap agent batches at a few hundred lines, require approval from a human who did not prompt the change, and make production unreachable from any agent credential, so a code freeze is a property of IAM rather than of a system prompt. None of that is new. That is the point.


Related: It's 2 AM. Do You Know What Your AI Agent Is Doing?, Agent Observability in 2026: The Spans Are Standard, the Standard Isn't Stable and CI/CD When Agents Open the Pull Requests.

I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.

Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.


Related

Oct 5, 2026

AIOps in 2026: 47% on ITBench, 10% on Hard On-Call, $6.50 an Investigation, and Four Outages Where the Automation Was the Incident

The incident-response agents of 2026 against the evidence: Datadog Bits AI SRE GA at about 6.5 credits ($6.50) per investigation, PagerDuty's approval-gated SRE Agent, Grafana's six GA agent tools, Splunk's AI SRE, Elastic buying Deductive, Resolve at a $1B headline valuation on $4M ARR, Traversal's $48M; benchmarks from 13.8% on ITBench in 2025 to 47% in May 2026, 20.7% exact root cause on OpenRCA 2.0, 10% on hard ORCA-Bench and 40% hallucinated causes; what the AWS, Azure, Cloudflare and Google postmortems say about automation; and why a human should hold the button.

14 min
Oct 5, 2026

CI/CD in 2026: Agents Opened 1M PRs in 5 Months, Bots Write 1 in 5 Reviews, and Most Agent PRs Get No Human Look

What the pipeline looks like when AI agents open the pull requests: GitHub's coding agent went GA on 25 September 2025 and opened over a million PRs in five months, Copilot code review passed 60 million reviews and one in five on GitHub, CodeRabbit has reviewed 13 million PRs on $88 million raised, and the first studies find most agent PRs get no human review attention. The gates that still hold: the assigner cannot approve, an extra approval for bot authors, path fences, signed commits, SLSA provenance, and a workflow YAML that gives agent PRs their own lane.

13 min
Oct 5, 2026

Coding Agents in 2026: Cursor at $2B Then Sold to SpaceX, Cognition at $1B, 58% on Terminal-Bench, and the Only RCT Still Says Slower

The coding-agent market as the primary sources report it in October 2026: Cursor from $500M to $1B to $2B in nine months and then acquired by SpaceX, Cognition from $73M to a $1B run-rate at a $48B valuation, Claude Code past $500M, Lovable at $13.3B, nearly 140,000 organisations on Copilot; four re-pricings from requests to tokens; Terminal-Bench 4.0 topping out at 58.2 percent for $3,267 a run; METR's redesigned study at minus 4 to minus 18 percent against 1.4 to 2x self-reports; 45 percent insecure samples, 1.7x more issues per AI pull request, and curl's bug bounty closed.

14 min