CI/CD in 2026: Agents Opened 1M PRs in 5 Months, Bots Write 1 in 5 Reviews, and Most Agent PRs Get No Human Look
My platform exists because agents write code nobody has read yet. PandaStack, which I build and run, boots a Firecracker microVM so an agent can execute whatever it just produced without touching anything that matters; that is the affiliation, disclosed. What changed this year is where the unread code ends up. Now the agent opens a pull request against the repository, a second agent reviews it, a third agent bumps the dependencies it touched, and the human in the loop is a notification badge. I spent the week reading what GitHub, the review and runner vendors and the first academic studies actually say about this. The shape: authoring industrialised in 2025, review is industrialising now, and the controls that hold are the boring ones, rulesets, required checks, signatures and provenance, applied in a separate lane for authors who are not people.
Who is opening the pull requests
GitHub's Copilot coding agent went generally available on 25 September 2025 for every paid Copilot plan. Mechanically it is a GitHub Actions job: assign an issue or type into the agents panel, and the agent gets an ephemeral Actions environment, a copilot/ branch and a draft PR. The docs cap a session at 59 minutes and one PR per task, billed as one premium request plus the Actions minutes it burns. The Octoverse 2025 report counts over a million PRs from that agent between May and September 2025, against 43.2 million PRs merged a month platform-wide, up 23 percent, and 986 million commits pushed in the year, up 25 percent. A million in five months is under half a percent of merges; the growth is the point.
Then GitHub opened the door. Agent HQ, announced at Universe on 28 October 2025, named Anthropic, OpenAI, Google, Cognition and xAI as agents that would run inside GitHub, and on 4 February 2026 Claude and Codex arrived in public preview for Copilot Pro+ and Enterprise, assignable to the same issue in parallel. Outside GitHub's plans, Anthropic's claude-code-action at v1, MIT licensed, runs on an @claude mention or a scheduled prompt with your own API or cloud credentials. OpenAI's codex-action@v1 runs codex exec inside the runner with a safety-strategy that drops sudo by default and a sandbox of read-only, workspace-write or danger-full-access. Cursor's cloud agents, formerly background agents, give each task its own Linux VM from a snapshot or a Dockerfile in .cursor/environment.json, start from an @cursor comment, and return a PR with screenshots and logs, billed at model API rates. Devin's Core plan is $20 a month plus $2.25 per ACU, Cognition's unit of roughly fifteen minutes of work.
To the pipeline every one of these is the same object: a non-human identity with write access to a branch, producing commits faster than any team has reviewed before. I covered the identity half in Who Is This Agent?.
Who is reviewing them
The review side's numbers are bigger than the authoring numbers. GitHub's Copilot code review went GA in April 2025 and by 5 March 2026 had done 60 million reviews, more than one in five on GitHub, with 12,000 organisations running it on every PR. GitHub's own quality figure: 71 percent of reviews surface "actionable feedback", averaging 5.1 comments; 29 percent say nothing. Each review costs 13 premium requests of a Pro plan's 300 a month, and since 1 June 2026 it also consumes Actions minutes, because the agentic version runs on Actions infrastructure. On 1 September 2026 GitHub let Copilot approve pull requests: every review now carries an approval assessment, and an administrator can enable, off by default and scoped to file paths, an approval that counts toward the required-approvals rule, dismissed on new commits like a human one. That is the most consequential setting in this post.
The independents raised on the bottleneck. CodeRabbit's $60 million Series B in September 2025, led by Scale with Nvidia's NVentures, took it to $88 million raised, 2 million repositories and 13 million PRs reviewed; its pricing runs $30, $60 and $108 per developer per month, free for public repositories. Greptile closed a $25 million Series A from Benchmark the same month and charges $30 a seat with 50 review credits included and $1 per extra credit. Graphite, the stacked-PR tool with its own reviewer, was bought by Cursor on 19 December 2025 and lists $20 and $40 per user annually, with "Cursor Cloud Agents are now in Graphite" across the top of the page: the company selling the agent that writes the PR now sells the tool that reviews it.
| Tool | Role in the PR | Price | Scale or usage | Source |
|---|---|---|---|---|
| Copilot coding agent | author | paid Copilot; 1 premium request per session plus Actions minutes; 59 min cap | 1M+ PRs, May to Sep 2025 | GitHub |
| Claude and Codex on Agent HQ | author | Copilot Pro+ ($39/mo) or Enterprise, public preview | launched 4 Feb 2026 | GitHub |
| claude-code-action v1 | author, reviewer | MIT; your API or cloud credentials | GitHub | |
| codex-action v1 | author, reviewer | your API key; drop-sudo default, sandbox modes | OpenAI | |
| Cursor cloud agents | author | Pro $20, Pro+ $60, Ultra $200, plus model API rates | owns Graphite since Dec 2025 | Cursor |
| Devin | author | $20/mo plus $2.25 per ACU (about 15 min) | TechCrunch | |
| Copilot code review | reviewer, approver (opt-in) | 13 premium requests per review plus Actions minutes | 60M reviews, over 1 in 5 on GitHub, 12,000 orgs | GitHub |
| CodeRabbit | reviewer | $30 / $60 / $108 per dev per month | 13M PRs, 2M repos, $88M raised | CodeRabbit |
| Greptile | reviewer | $30 per seat, 50 credits, $1 per extra | $30M raised | Greptile |
| Graphite (Cursor) | reviewer, stacking | $20 / $40 per user, annual | acquired 19 Dec 2025 | Graphite |
What the evidence says about review quality
CodeRabbit's own report of December 2025, 320 AI co-authored against 150 human PRs on open source, found 10.83 issues per AI PR against 6.45, the 1.7x headline, with logic errors up 75 percent and security findings up to 2.74x. It is a vendor measuring with its own reviewer on a sample labelled by inference, and it says so; a direction, not a rate.
The academic work is more careful and more worrying. An EASE 2026 paper from Nicolaus Copernicus University on the AIDev dataset found that most AI-generated PRs receive no review attention at all, and that when reviewed the review is "automation-mediated": a human steering the agent in comments rather than evaluating the change. Human PRs in the same repositories were far more likely to get a human-only review; the authors' conclusion, which I would underline for anyone building dashboards, is that review counts no longer measure human oversight. An ESEM 2026 study of 248,641 agent-authored PRs that received an agent review found cross-vendor review (CodeRabbit on a Claude Code PR, say) in only 1.6 percent, though that volume grew by more than two orders of magnitude across 2025; CodeRabbit labelled 35 percent of its comments on Claude Code PRs as refactors against 10.5 percent on Copilot PRs; agents reviewing agents have house styles. And the longitudinal study of one company's "2x mandate", 802 developers and 196,212 PRs from January 2024 to April 2026, reached 2.09x the pre-mandate throughput per developer with merge and revert rates steady, while reviewer load roughly doubled and "automated review overtook human review". Nothing broke in the data; the humans just stopped being the reviewers.
Flaky tests are the quiet multiplier. Google's figures, old but the best measured, are that almost 16 percent of its tests show some flakiness and about 84 percent of pass-to-fail transitions in CI involve a flaky test. An agent told to make CI green will retry, then "fix" the test, then skip it, and a reviewer who sees a green check and an approval assessment will merge. I have found no measurement of how often that happens; I assume it happens wherever the agent can edit tests and the human does not read them.
The CI bill
Agents spend CI minutes twice: once to write the change and once to check it. GitHub cut hosted-runner prices on 1 January 2026 by up to 39 percent, Linux 2-core from $0.008 to $0.006 a minute inclusive of a new $0.002 platform charge, and postponed, after a loud December, the same charge on self-hosted runners in private repositories; it says 96 percent of customers saw no bill change. The runner vendors sell the gap between a stock VM and a tuned one. Blacksmith prices Ubuntu x64 at $0.004 a minute and ARM at $0.0025, claims 2x hardware and 4x cache throughput, and so claims a 67 percent saving. Depot charges the same $0.006 as GitHub but bills per second and publishes its wins: grpc 7.0x faster (25m 26s to 3m 38s), Kafka 2.7x, Zed 2.4x. Namespace sells unit-minutes on its own AMD EPYC and Apple M5 Max hardware and cites 4x at Warp. None of these claims is independent and all are plausible. An agent that runs 59 minutes on a slow runner costs an hour of minutes and an hour of wall clock, and the second number is the one reviewers feel. The cost not on the invoice is review attention, and Copilot code review consuming Actions minutes from June is GitHub saying out loud that review is now compute.
Gates that still hold for a bot author
GitHub's own mitigations for the Copilot agent are a template for any agent, because the platform enforces them, not the model: only users with write access can trigger it; it can push only to its own PR branch or a new copilot/ branch; its PRs are drafts it cannot mark ready, approve or merge; the person who assigned the task cannot count as the approver; under rulesets an agent PR needs one more approval than a human one, on by default; and workflows do not run on its commits until a human with write access clicks "Approve and run workflows". Internet access is restricted, hidden Unicode is stripped from prompts, and its commits are signed.
The same shape works for everyone else if you build it. Dependabot's docs are the oldest version: branch on github.event.pull_request.user.login == 'dependabot[bot]', read dependabot/fetch-metadata, and gh pr merge --auto only for update types you trust, with required status checks on so auto-merge cannot outrun CI; Renovate has had the equivalent for years. CODEOWNERS still routes the files that matter to their owners, and a code-owner review requirement is satisfied only by that owner, never by a bot, unless an administrator enables Copilot approvals for that path. My rule: bot identities never go on a bypass list and never satisfy a required review on .github/, deploy or infrastructure paths. Here is the workflow I would run.
# .github/workflows/pr-gate.yml
# One workflow, two lanes. Human-authored PRs run the normal checks.
# PRs opened by an agent identity (any GitHub App bot such as
# copilot-swe-agent[bot], dependabot[bot] or renovate[bot], or a branch
# named copilot/*, codex/*, cursor/*) get a stricter lane: superseded runs
# are cancelled, there is a hard time budget, pipeline and deploy paths are
# fenced, every commit must be signed, no bot may certify the DCO, and the
# build artifact gets SLSA provenance. Mark the final `gate` job as the one
# required status check in your ruleset and keep bots off the bypass list.
name: pr-gate
on:
pull_request:
types: [opened, synchronize, reopened, ready_for_review]
permissions:
contents: read
# Agents push commits in bursts; keep only the newest run per PR.
concurrency:
group: pr-gate-${{ github.event.pull_request.number }}
cancel-in-progress: true
env:
BASE: ${{ github.event.pull_request.base.sha }}
HEAD: ${{ github.event.pull_request.head.sha }}
jobs:
classify:
runs-on: ubuntu-24.04
outputs:
agent: ${{ steps.c.outputs.agent }}
steps:
- id: c
env:
AUTHOR: ${{ github.event.pull_request.user.login }}
AUTHOR_TYPE: ${{ github.event.pull_request.user.type }} # "User" or "Bot"
BRANCH: ${{ github.head_ref }}
run: |
agent=false
[ "$AUTHOR_TYPE" = "Bot" ] && agent=true
case "$BRANCH" in copilot/*|codex/*|cursor/*) agent=true ;; esac
echo "author=$AUTHOR type=$AUTHOR_TYPE branch=$BRANCH agent=$agent"
echo "agent=$agent" >> "$GITHUB_OUTPUT"
agent-fence:
needs: classify
if: needs.classify.outputs.agent == 'true'
runs-on: ubuntu-24.04
timeout-minutes: 10
steps:
- uses: actions/checkout@v5 # pin to a commit SHA in your own repo
with:
fetch-depth: 0
- name: Agents may not change the pipeline, ownership or deploy config
run: |
if git diff --name-only "$BASE" "$HEAD" | grep -E '^(\.github/|CODEOWNERS$|deploy/|terraform/)'; then
echo "::error::agent-authored PR touches protected paths; a human must open this change"
exit 1
fi
- name: No bot may certify the Developer Certificate of Origin
run: |
if git log --format=%B "$BASE..$HEAD" | grep -Ei '^Signed-off-by:.*(\[bot\]|copilot|codex|cursor|devin)'; then
echo "::error::Signed-off-by from a non-human identity"
exit 1
fi
- name: Every commit in the PR must carry a GitHub-verified signature
env:
GH_TOKEN: ${{ github.token }}
run: |
for sha in $(git rev-list "$BASE..$HEAD"); do
v=$(gh api "repos/$GITHUB_REPOSITORY/commits/$sha" --jq '.commit.verification.verified')
[ "$v" = "true" ] || { echo "::error::commit $sha is unsigned or unverified"; exit 1; }
done
test:
needs: [classify, agent-fence]
# Humans skip the fence; agents must have passed it.
if: always() && needs.classify.result == 'success' && needs.agent-fence.result != 'failure'
runs-on: ubuntu-24.04
timeout-minutes: 20 # the upstream agent session is capped at 59 min; your tests should not be
steps:
- uses: actions/checkout@v5
with:
fetch-depth: 0
- run: make test
- name: Agents do not get to skip or disable tests quietly
if: needs.classify.outputs.agent == 'true'
run: |
if git diff "$BASE" "$HEAD" -- 'tests/**' '**/*_test.*' '**/*.test.*' | grep -E '^\+.*(skip|xfail|\.only\(|t\.Skip)'; then
echo "::error::agent PR adds a skip/xfail/only marker; read the test change before the code"
exit 1
fi
provenance:
needs: [classify, test]
if: needs.classify.outputs.agent == 'true' && needs.test.result == 'success'
runs-on: ubuntu-24.04
permissions:
contents: read
id-token: write # Sigstore keyless signing
attestations: write
steps:
- uses: actions/checkout@v5
- run: make dist # writes dist/app.tar.gz
- uses: actions/attest-build-provenance@v4 # SLSA v1.0 Build L2 here; L3 needs a reusable workflow
with:
subject-path: dist/app.tar.gz
gate:
# The only job you mark "required" in the ruleset. Skipped lanes are fine; failed or cancelled are not.
needs: [classify, agent-fence, test, provenance]
if: always()
runs-on: ubuntu-24.04
env:
NEEDS: ${{ toJSON(needs) }}
steps:
- run: |
if echo "$NEEDS" | grep -Eq '"result": "(failure|cancelled)"'; then
echo "::error::a lane failed"; exit 1
fi
echo "all lanes green"
Mark gate as the single required status, keep GitHub's extra-approval rule for agent PRs on, and leave Copilot's countable approvals off except, if you must, on docs and fixtures. The path fence is not there because agents write bad Terraform. A workflow file is the one place a PR can change what the pipeline itself does, which is exactly the move an injected agent makes; the npm worms of the past year went straight for CI credentials once they had an agent's hands.
Provenance for commits nobody typed
Who committed this is now a question with a cryptographic answer. Sigstore's gitsign, v0.17.1 since August, signs commits keylessly: an OIDC identity gets a ten-minute Fulcio certificate and the signature lands in the Rekor log, so an agent under a workload identity in Actions leaves a signature naming the workflow, repository and run. GitHub does not show gitsign signatures as "Verified", because Sigstore's root is not in its trust store and the certificate has expired by the time anyone looks; verification means gitsign verify against Rekor, not a green badge. For the artifact, GitHub's artifact attestations, via actions/attest-build-provenance at v4.2.2, give SLSA v1.0 Build Level 2 by default and Level 3 from a reusable workflow, checked with gh attestation verify. That ties an artifact to a commit, a workflow and a run; it does not say whether a model wrote the commit.
For that, the convention that is winning is a trailer. The Linux kernel's coding-assistants document specifies Assisted-by: LLM [TOOL1] [TOOL2] and the rule that matters: "AI agents MUST NOT add Signed-off-by tags. Only humans can legally certify the Developer Certificate of Origin." Red Hat's emerging-technologies group proposes the same two-signature pattern for agents themselves, Sigstore at build time and SPIFFE at runtime. The step in my workflow that rejects a bot Signed-off-by is three lines; it protects the one legal statement in the pipeline that must come from a person.
The projects that said no, and the one that said yes
The clearest signal of how maintainers feel is what they have banned. curl ended its bug bounty on 31 January 2026 after six years, $86,000 and 78 confirmed vulnerabilities, because about 20 percent of reports had become AI-generated and the valid rate had fallen to 5 percent; real reports still go through GitHub Security Advisories. Gentoo's council voted on 14 April 2024 that contributing content created with NLP AI tools is "expressly forbidden", on copyright, quality and ethical grounds. NetBSD's commit guidelines presume LLM output "tainted" and bar it without core's written approval. QEMU's code-provenance policy is to "DECLINE any contributions which are believed to include or derive from AI generated content", since nobody can honestly sign the DCO for it. Zig's code of conduct goes furthest: no LLM-generated code or prose, no paraphrasing it, no LLM editing, translation or bug-finding, and no discussing chatbot use, because the project spends review time to grow contributors and an LLM-assisted PR returns nothing on that investment.
Debian went the other way. Its general resolution closed on 29 August 2026 with "Responsible Use of Generative AI" winning: Debian "neither endorses nor prohibits" the tools, every contribution must meet the same standards however produced, disclosure is encouraged not required, and both ban options lost to "none of the above". The OpenSSF and CNCF published Securing Open Source in the Age of AI on 19 May 2026 for maintainers handling AI-generated contributions and reports, naming hallucinations, slopsquatting and inflated severity scores as the failure modes. Both responses are rational. A volunteer project with no review budget bans; a project with a review process and a DCO raises the bar for the human who signs.
What I take from this. The agents are not the risk; the merge button is. PR volume up 23 percent, a million agent PRs in five months, one review in five by a model, most agent PRs never read by a person, and a GitHub setting, off today, that lets the model's approval count. The discipline this calls for is the part of DevOps that was always about trust boundaries, not speed: who may push where, what must pass, who must sign, what the artifact proves. This quarter, on my own repositories: classify agent PRs as a lane with the workflow above and make gate the only required check; keep every bot off the bypass list and leave Copilot approvals disabled; require signed commits and reject any Signed-off-by from a bot; and read the tests an agent changes before its code, because the test is where a tired reviewer and an eager agent agree to lie to each other. The sandbox was never the hard part. The hard part is the moment the code leaves it, the moment I described in It's 2 AM. Do You Know What Your AI Agent Is Doing?, and that moment now has a pull request number.
Related: Who Is This Agent? Non-Human Identity in 2026, The Worms Learned to Use Your AI Agent: A Year of npm Supply-Chain Attacks and MCP Security in 2026: The Protocol Got Hardened. The Ecosystem Didn't..
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
CRA Reporting Went Live on 11 September: 24 Hours to ENISA, 3.08 Billion Rekor Entries, 17% of PyPI Attested, 92% SBOM False Positives
Supply-chain compliance in 2026: what the Cyber Resilience Act's Article 14 duty requires since 11 September 2026 (24-hour early warning, 72-hour notification, ENISA's Single Reporting Platform), what waits until 11 December 2027, how the open-source steward role works, where SBOMs stand (CISA's 2026 minimum elements, CycloneDX 1.7, a 92 percent false-positive rate in a 2,414-repo study), how far provenance has got (SLSA 1.2, 20 percent of PyPI uploads via trusted publishing, Rekor at 3.08 billion entries), the US retreat from mandates, and the pipeline I would run.
13 minSep 8, 2026AI Found the Bugs: 23,000 Findings in a Month, a 3-Day Patch Mandate, and the End of curl's Bounty
From AIxCC's 18 real bugs in August 2025 to Project Glasswing's 23,019 findings by May 2026 and the first AI-written zero-day exploited in the wild. What the cyber reasoning systems, Big Sleep, Codex Security and Mythos Preview actually did, why bug bounties are drowning, why CISA now wants federal patches in three days, and a workflow for a small team that keeps the humans looking only at validated findings.
13 minOct 5, 2026AI SOC Agents in 2026: 98% Accuracy Claims, 23 to 34% on the Benchmark, and 0% of Teams Letting Them Act Alone
What the security-operations agents from Microsoft, Google, CrowdStrike, Palo Alto, SentinelOne, Torq and a billion-dollar startup cohort actually do in 2026 and what is measured: a median of 100 alerts a day and 28 percent never investigated, 75-minute mean investigations, Microsoft's $4-an-hour compute units and Google's token meter, CrowdStrike's 98 percent triage claim against Meta's 23 to 34 percent benchmark, Anthropic's and Google's reports of attackers running agents against defenders, and a Sigma rule and CI step that keep a human on the merge button.
14 min