Confidential Computing for AI in 2026: Apple Moved to Google's TEEs, a $1,000 Interposer Broke the Keys, and Firecracker Still Can't Do It

Sep 8, 2026 · 15 min · Ajay Kumar

The question I get most from enterprise prospects, after "how fast does it boot", is whether their data can be protected from me. Not from other tenants; I have written enough about the kernel boundary between microVMs. From the operator. Can the host, the hypervisor, the person with root on the machine, be prevented from reading what the agent is doing? The honest answer for my platform today is no, and this post is the long form of why, what would change it, who has done it for AI inference, and what the guarantee is actually worth after the year the hardware had.

What a TEE promises, precisely

A confidential VM runs with its memory encrypted by a key the CPU holds and the host never sees. AMD calls it SEV-SNP, Intel calls it TDX, Arm's version is CCA and is not shipping yet. The host can still schedule the VM, starve it, kill it and watch its network traffic, but it cannot read its registers or memory, and the CPU will sign a report saying what firmware and initial image the VM was launched with. That signed report is the attestation, and it is the whole product: a remote party can check the signature against the vendor's certificate chain, check the measurement against the image they expect, and only then release a key or a secret to the VM. RFC 9334 defines the roles: the attester produces evidence, a verifier appraises it, a relying party acts on the result.

For AI the CPU is not enough, because the model runs on a GPU, and the GPU has its own memory over its own bus. NVIDIA's Hopper generation added a confidential mode where traffic across PCIe is encrypted through bounce buffers and the GPU produces its own attestation, verified through NVIDIA's remote attestation service. Blackwell adds encrypted NVLink and is the first GPU marketed as capable of TEE-I/O, the standard for letting a device join the CPU's trust boundary directly. Intel's Trust Authority will now do a composite check, a TDX quote plus up to eight GPU reports, and hand back one signed token with a tdx section and an nvgpu section. That token is what "confidential inference" means when a vendor says it.

Host machine: hypervisor, operator with root, other tenants sees: scheduling, memory size, network metadata, power. Cannot read guest memory. trust boundary (hardware-enforced) Confidential VM SEV-SNP or TDX memory encrypted, CPU-held key agent runtime, tokenizer, prompt in plaintext only here launch measurement signed GPU in CC mode H100 / H200 / Blackwell PCIe bounce buffers encrypted NVLink encrypted (Blackwell) weights + KV cache in HBM GPU attestation report Remote verifier (customer, or Intel Trust Authority, Azure MAA, Trustee) 1. AMD/Intel cert chain OK? 2. NVIDIA NRAS: GPU report OK? 3. measurement = expected image? 4. TCB version not revoked? → token → release keys reports keys Outside the guarantee: availability (host can stop or starve you) · side channels (cache, timing, perf counters, power) · I/O and traffic metadata physical bus attacks (TEE.fail, Battering RAM, WireTap: an interposer on the DIMM recovers the attestation keys) bugs in the code you attested · prompt injection inside the agent · the vendor's firmware, which is in your TCB
What a confidential inference deployment looks like and, in the lower strip, the list of things the signed report does not cover. The strip is the part to read twice.

What it costs

The overhead is mostly the encrypted transfers between host and GPU, so it depends heavily on what you measure. The published numbers, all on H100:

Study Setup Overhead
Phala et al., 2024 H100 CC, LLM queries below 7%, near zero for large models and long sequences
ETH Zurich, Sep 2025 H100 CC; TDX and SGX on CPU 4 to 8% GPU throughput; CPU TEEs under 10% throughput, under 20% latency
Mozilla, 2026 GCP a3-highgpu-1g, TDX + H100, vLLM time-to-first-token +21.8% (Mistral 7B) and +27.8% (Qwen3 30B-A3B); throughput −18 to −21% at fixed rate; 11 to 20% at saturation
IBM, 2024 vLLM in CPU and GPU enclaves "negligible" with parallelism

NVIDIA's own page says "up to 98% of the performance of unsecured deployments." Mozilla's measurement is the one to plan on, because it was taken on a production cloud instance with a production serving engine rather than a lab, and its advice is to provision 15 to 25% extra capacity. The pattern is consistent with the mechanism: steady-state decode barely notices, prefill and small-batch latency notice a lot, because that is where the bytes cross the bus. Read alongside my inference post: if you have disaggregated prefill from decode, confidential mode taxes the prefill pool.

Where you can get it

Not from AWS, which is the interesting absence. AWS's confidential computing page argues that the Nitro system itself gives operators "no mechanism" to access instance memory, and at re:Invent 2025 it launched instance attestation, NitroTPM-signed documents and KMS policies keyed on boot measurements, on GPU instances included, plus a formally verified isolation engine on Graviton5. That is a strong operational and design claim. It is not a hardware TEE with a vendor-signed memory-encryption guarantee, and no AWS instance I can find runs NVIDIA GPUs in confidential mode. SEV-SNP is available on a few AMD instance families in two regions, per third-party comparisons I could not confirm on AWS's own docs.

Azure has the longest list: SEV-SNP families on Milan and Genoa, TDX families that went GA this February, and the only confidential GPU SKU on the platform, NCC H100 v5, one H100 with 94 GB per VM, GA since September 2024 in two regions. Nothing Blackwell yet.

Google moved fastest this year. Its release notes have the TDX-plus-H100 instance GA since July 2025, and on 24 June this year it announced Confidential G4, AMD Turin with RTX PRO 6000 Blackwell GPUs, in preview, alongside TDX on Granite Rapids, live migration for confidential VMs, GA for Confidential Space with Hopper GPUs and Intel's attestation service, and an open-source SDK for encrypting prompts end to end. There is also a live advisory that SEV-SNP instances may see longer boots and "performance changes" from August to November because of a guest-kernel migration, which is the sort of detail that reminds you a confidential VM is still a VM.

And the reason Google moved fastest is Apple. On 8 June Apple announced that Private Cloud Compute, the inference system it built in 2024 on its own silicon with a public transparency log so that researchers can verify what software is running, now also runs on Google Cloud: Intel TDX CPUs, NVIDIA Blackwell GPUs in confidential mode, Google's Titan chip as a root of trust, an append-only ledger of the fleet's hardware, and attestation "rooted in at least two separate roots of trust from independent vendors." It is in preview and "gradually ramping towards the complete set of protections." That last phrase is doing work. Neither Apple nor Google says whether the GPUs are attached with TEE-I/O or the older bounce-buffer mode, and I could find no confirmation that Intel's TDX Connect, the CPU side of TEE-I/O, is in production on any cloud; Intel's own blog last year said the host and device ecosystem would be ready "later in 2025." I would read Apple's move as the strongest endorsement the technology has had, and the "gradually ramping" as an accurate description of where TEE-I/O is.

The others: Google's own Private AI Compute runs Gemini on TPUs inside what it calls Titanium Intelligence Enclaves with an IP-blinding relay, and NCC Group's assessment of it found a timing side channel in the relay and three attestation-protocol issues. Meta's Private Processing for WhatsApp uses confidential VMs and confidential-mode GPUs behind an oblivious HTTP relay, with promised binary logs and a bounty. A cluster of startups sell the same thing as a service: Edgeless Systems' Privatemode on Blackwell, which cites a German electronic-health-record deployment; Tinfoil, with containers on TDX or SEV-SNP and up to eight H200 or B200 with a Sigstore transparency log; Phala at $3.08 an hour per H100 with dual attestation; Fortanix; NEAR.

And then the two labs everyone actually sends prompts to. OpenAI's August post on data confidentiality reaffirms zero data retention, previews something called Private Safety Processing and promises a whitepaper in September; it makes no attestation claim. Anthropic's only public artefact is a June 2025 whitepaper with Irregular that distinguishes bridged designs, where a CPU enclave mediates a non-confidential accelerator, from native accelerator TEEs, and sets a target of the RAND report's highest security levels; I found no announcement of a production confidential inference product. Both companies' confidentiality story to enterprises today is contractual and operational, not attested. That is the same as mine, which is some comfort and no answer.

The year the hardware got embarrassed

If the guarantee is a signed report, the attack is to sign your own. Three academic groups did exactly that in the last twelve months with hardware costing less than a GPU-hour.

Battering RAM, from KU Leuven and Birmingham, is a $50 interposer on a DDR4 DIMM that breaks SGX and, for SEV-SNP, replays a launch measurement so the attestation is forged; advisories were published on 30 September 2025. WireTap, from Georgia Tech and Purdue, extracts SGX attestation keys with a sub-$1,000 DDR4 interposer. And TEE.fail, from the same group, does it on DDR5: a sub-$1,000 interposer that recovers SGX and TDX attestation keys and the SEV-SNP signing keys on Zen 4 and Zen 5, and thereby breaks NVIDIA's confidential mode too, because the GPU trusts the CPU's forged attestation. Intel's advisory is dated 28 October 2025 and AMD's is SB-3040. The vendors' position, uniformly, is that physical attacks are outside the threat model. That is defensible for a chip vendor. It is a problem for a customer whose threat model is "the cloud operator," who has physical access by definition; the mitigation is that the operator is a large company with audits, which was the guarantee before TEEs existed.

The software side was busy too: BadRAM, a sub-$10 modification to a DIMM's configuration chip that AMD fixed in firmware; the interrupt-injection attacks Heckler and WeSee; TDXdown's single-stepping; CounterSEVeillance, which recovered an RSA key from performance counters that AMD confirmed are "not protected" until Zen 5; and BadAML at CCS 2025, which compromised SNP VMs through unattested legacy firmware interfaces, a reminder that the launch measurement covers the initial image and nothing that runs afterwards unless you chain it into a vTPM. On that chaining: on Azure the SNP report sits at TPM NV index 0x01400001 and measures Microsoft's paravisor, which hosts the vTPM, so the paravisor is in your trusted computing base and you are trusting Microsoft's measurement of Microsoft's code. Every cloud has an equivalent. "Verify, don't trust" is the slogan; in practice you verify a chain that terminates in the cloud vendor's firmware and the chip vendor's key, and you choose whom to trust about the DIMM.

The kernel and the microVMs

For the operator rather than the customer, the question is what the software stack can do. The kernel status as of this month: SEV-SNP guest and host support have been upstream for a while; Intel's TDX host side merged in 6.16 in June 2025 after years out of tree, along with the SVSM virtual TPM driver that lets a confidential guest have a measured boot chain without trusting the host; Arm CCA's guest side merged in 6.14 but the host side was at version 14 of the patch series in May with acknowledged gaps, and I could not find a generally available CCA instance anywhere. RISC-V's spec has not been ratified since a 2024 draft.

Above the kernel, Cloud Hypervisor supports SEV-SNP and TDX via IGVM images, mutually exclusive, with no hotplug, and merged ID-block support in May. The Confidential Containers project shipped v0.22 in July with its Trustee attestation stack and Kata 4.1 in August; it is the Kubernetes answer and it works, at the cost of each pod being a full confidential VM.

Firecracker, which my platform runs on, has none of it. There are no SEV-SNP or TDX pull requests or open issues in the repository as of this morning; the 2020 request was closed. What there is, on a branch last touched on 14 August, is "secret hiding", the kernel's guest_memfd work that unmaps guest memory from the host's page tables so a host compromise cannot casually read it. That is a real hardening and I want it, but it is not encryption and it does not attest. The technical reason a microVM monitor is slow to get TEEs is that everything that makes Firecracker fast, the snapshot-restore boot I have measured to the millisecond, the copy-on-write memory forks, the userfaultfd demand paging, requires the host to read and write guest memory, and a TEE's entire purpose is to stop that. Restoring a 179 ms snapshot into an encrypted guest means re-encrypting under a fresh key and re-measuring, and the launch flow for an SNP guest is seconds, not milliseconds. Confidential microVMs will exist; they will not boot the way mine do, and a platform that offers both will be running two very different fleets.

Checking it yourself

If a vendor tells you a workload is confidential, ask for the report and check it. On an Azure SEV-SNP VM, from Microsoft's own guide:

curl -H Metadata:true http://169.254.169.254/metadata/THIM/amd/certification > vcek
jq -r '.vcekCert , .certificateChain' vcek > vcek.pem
sudo tpm2_nvread -C o 0x01400001 > snp_report.bin
dd skip=32 bs=1 count=1184 if=snp_report.bin of=guest_report.bin   # strip Azure's 32-byte header
snpguest verify attestation . guest_report.bin                      # "VEK signed the Attestation Report!"

On any SNP guest with snpguest 0.10: snpguest report report.bin request.bin --random, then fetch the certificate chain for your CPU generation, snpguest verify certs, snpguest verify attestation, and snpguest display report to read the measurement and TCB versions you then compare against what you expected. For the GPU, on GCP or Azure:

sudo nvidia-smi conf-compute -f     # CC status: ON
sudo nvidia-smi conf-compute -grs   # Confidential Compute GPUs Ready state: ready

Then NVIDIA's nvtrust for the GPU report, noting that the Python SDK is deprecated in favour of a C++ CLI as of June, and that the multi-GPU "protected PCIe" mode leaves NVLink traffic in plaintext inside the boundary. A signed report that you did not verify against a measurement you computed is a PDF.

My position, then, for the prospects who ask. Confidential inference is real, it is on Google and Azure with a 15 to 25% latency tax, Apple has staked its brand on it, and the attestation tooling is good enough that a customer can check it without trusting the operator's word. It does not protect against a physical attacker, which the cloud operator is; that protection is still contractual. The labs you actually call do not offer it yet, so the prompt is in plaintext somewhere regardless. And on a microVM platform built for 179 ms boots, it is not available, which I would rather say than imply. When a Firecracker-shaped monitor can launch an SNP guest from a snapshot in under a second, I will build the second fleet.


Related: Container vs microVM vs gVisor: Choosing Agent Isolation, I Measured Every Stage of My 179ms Firecracker Boot Path and Post-Quantum Migration in 2026.

I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.

Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.


Related