AI for Science in 2026: One Phase III Trial, Two Operational Weather Models, and Erdős Proofs Nobody Wants to Read

Sep 9, 2026 · 16 min · Ajay Kumar

I am not a scientist. I run a microVM fleet, and the reason I spent a week reading AI-for-science papers is that "AI for science" is now a compute customer: the US Department of Energy is building a 100,000-GPU machine for it, three startups raised more than a billion dollars between them on the pitch, and my own prospects have started asking whether an agent can drive a protein-design pipeline inside a sandbox. So I wanted to know what is real. Not what the press releases say, which is uniformly "breakthrough", but what has been checked by someone other than the people who built the model. This is that list, sorted by how much independent validation exists, with the pip-install status of everything I could confirm.

The ladder

Every claim in this field sits on one of four rungs. A benchmark number, usually on the vendor's own benchmark. Open code and weights, which lets someone else run it. Independent replication, where a lab that did not build the model confirms the result. And operational or clinical use, where a forecaster issues warnings from it or a regulator lets patients take it. Most of the money is on rung one. Almost all of the durable value is on rungs three and four. The figure is where I put each domain after the reading.

Where each domain stands, September 2026 Proteins Materials Weather Mathematics Agentic science Operational /clinical use Independentreplication Open codeand weights Vendorbenchmark rentosertib Phase III started 7 Jul 2026 no approvals yet none found ECMWF AIFS Feb/Jul 2025 NOAA AIGFS Jan 2026 NHC used cyclone model IMO-graded gold 2025 disproof verified by Alon, Gowers et al. none Evo 2: 16/285 phage genomes viable; metallo- hydrolases in Nature critiques instead: GNoME "scant evidence" MatterGen 1972 compound ECMWF's own scorecards: worse on 10 m wind, under-predicts extremes Tao's wiki: many "solves" are literature rediscoveries Co-scientist: 6 external labs (Nature, May 2026) Kosmos 79% accurate AF3 weights academic only Boltz-2 MIT, OpenFold3 Apache; Chai-2 closed MatterGen MIT, UMA open, Orb Apache, OMol25 100M+ DFT calculations AIFS via anemoi, Aurora MIT, WeatherNext Cyclones Apache Deep Think, Aletheia, OpenAI's model: closed; Lean proofs are public Kosmos and Sakana code public; hosted models behind them IsoDDE "2x AF3", Chai-2 "86% developable", Latent-X "90% hit rate" (vendor) GNoME 2.2M crystals, Periodic Labs $300M seed, no results published GenCast beats ENS on 97.2% of targets; Aurora 1.5 on 88.9% (vendor) FrontierMath v2, HLE 46.5% top score, IMO 2026 claims unverified LABBench2: frontier models drop 26 to 46% vs LAB-Bench v1 green: reached with evidence I could read · tan: reached partially or with caveats · grey: not reached
Five domains against four rungs of validation. The weather column is the one I would show a sceptic; the materials column is the one I would show an investor.

Weather: the one that is actually in production

Start with the domain where the argument is over. On 25 February 2025 ECMWF put its AIFS Single model into operational use, and on 1 July 2025 it followed with AIFS ENS, a 51-member ensemble at roughly 31 km with 229 million parameters. ECMWF is the most conservative forecasting centre on earth and it now issues AI forecasts alongside its physics model, at what it says is around a thousandth of the energy per forecast. NOAA followed on 5 January 2026 with AIGFS, AIGEFS and HGEFS, and the reporting I saw says AIGFS is a GraphCast derivative fine-tuned on NOAA's own analyses using a fraction of a percent of the GFS compute. Google's cyclone model was used by the National Hurricane Center during the 2025 season, and Google open-sourced it on 6 August this year.

What makes this column trustworthy is not the vendor scores, although those are striking: GenCast beat ECMWF's ensemble on 97.2 percent of 1,320 targets in the Nature paper, and Microsoft says Aurora 1.5 beats it on 88.9 percent. What makes it trustworthy is that ECMWF publishes its own scorecards against itself, and they are not flattering everywhere. Newsletter 185 says AIFS ENS underperforms the physics ensemble on 10 m wind, under-predicts extreme precipitation and is over-dispersive aloft. A model whose operator tells you where it is worse is a model you can plan around. Google's WeatherNext 3, announced on 3 September at 5 km hourly with precipitation claims of 60 percent better than satellite estimates, is being scored live by an independent third party, which is the right way to make that claim.

For an engineer, the whole stack is runnable. AIFS runs through the anemoi-inference package, which had a release on 7 September; the older ai-models package is stale since December 2024, so do not start there. Aurora is on PyPI as microsoft-aurora 2.0.1 under MIT. NVIDIA's Earth-2 open model family launched in January. The compute is small: these are hundreds of millions of parameters, inference on a single GPU in minutes, which is why national centres could adopt them in a year rather than a decade.

Proteins: one drug in Phase III and a licensing map

The structure-prediction story has moved from "can we predict the fold" to "who is allowed to use the predictor". AlphaFold 3's code has been public since November 2024 under CC-BY-NC-SA, and the weights are academic-only on request. Nothing changed in 2025 or 2026. The open reimplementations are where the commercial world lives:

Model Package Version, date Licence Notes
Boltz-2 boltz 2.2.1, 8 Sep 2025 MIT structure plus binding affinity; no release in 12 months
OpenFold3 openfold3 0.5.0, 21 Aug 2026 Apache-2.0 only AF3-class model with permissive code and weights
Chai-1 chai_lab 0.6.1, 18 Mar 2025 Apache-2.0 Chai-2 is not released; preprint only
ESM esm 3.4.0, 27 Aug 2026 MIT code check per-weight licences; ESM3 weights were non-commercial
Evo 2 evo2 0.6.0, 19 Jun 2026 Apache-2.0 7B and 40B genome models, Nature 4 Mar 2026
RFdiffusion3 GitHub 3 Dec 2025 see repo "ten-fold faster" than RFdiffusion2

I checked each of those version dates against the PyPI JSON API this morning; the licences I read from the repositories. The pattern I want to flag is the same one I found in open-weight LLMs: the headline model is restricted, the reimplementation is permissive, and the reimplementation is a year behind. OpenFold3's default weights have a June 2025 training cutoff.

On the replication rung, two results stand out because a wet lab confirmed them. Arc Institute's Evo 2 designed 285 phage genomes and 16 propagated in the lab, a 5.6 percent hit rate on an absurdly hard task. And the Baker lab's metallohydrolases designed with RFdiffusion2 were published in Nature on the same day RFdiffusion3 shipped. Company benchmarks are a different rung: Isomorphic's IsoDDE report claims more than twice AF3's accuracy on a protein-ligand generalisation benchmark that Isomorphic wrote, Chai-2's antibody preprint says over 86 percent of designed antibodies have strong developability profiles, and Latent-X claims over 90 percent hit rates on macrocycles. All of those may be true. None has been reproduced by anyone without a stake in it.

The clinical rung has exactly one entry that matters. Insilico's rentosertib, a TNIK inhibitor for idiopathic pulmonary fibrosis whose target and molecule were both AI-generated, published a Phase IIa in Nature Medicine in June 2025: 71 patients, 22 Chinese sites, 12 weeks, a 98 mL forced-vital-capacity gain in the 60 mg arm against a 20 mL loss on placebo. On 7 July 2026 it entered Phase III, 320 patients across 47 centres with a 52-week endpoint. That is a small, short, single-country Phase IIa followed by a real Phase III. It is also the entire pipeline of AI-discovered drugs at that stage. Recursion's REC-4881 reports 43 to 53 percent polyp-burden reduction in Phase 1b/2; Xaira, which raised a billion dollars in 2024, has disclosed no clinical candidate; and Isomorphic, which closed a $2.1 billion Series B on 12 May 2026, names no trial date in the release. Its CEO's "first-in-human by end of 2026" appears only in trade press, and I could not verify it.

Materials: the column with critiques where results should be

The materials column is the one where I changed my mind. DeepMind's GNoME announced 2.2 million new crystals in 2023, and in April 2024 Cheetham and Seshadri published a Chemistry of Materials analysis finding "scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility." The autonomous A-Lab that was meant to synthesise them was critiqued by Palgrave and Schoop, who found around two-thirds of the "new" compounds were ordered versions of known disordered phases. I found no 2025 or 2026 DeepMind paper reporting GNoME materials made and validated. Microsoft's MatterGen, published in Nature in January 2025, has a showcase compound, TaCr₂O₆, whose measured bulk modulus came in at 169 GPa against a 200 GPa target; a 2026 Materials Horizons paper argues it is identical to a compound from 1972 that was in the training set. I could only read that critique through snippets, so treat my summary as second-hand.

What is real here is the open data and the interatomic potentials. Meta's OMol25 is over 100 million density-functional-theory calculations, six billion core-hours of compute given away, and the UMA models trained on it are open. Orbital's Orb-v3 is Apache-2.0 on PyPI. These are force fields, not discovery engines, and they make molecular simulation cheaper for everyone. The discovery engines are where the money went: Periodic Labs raised a $300 million seed, Lila Sciences $550 million total, CuspAI reportedly $450 million in July. I could find no published scientific result from Periodic Labs at all; the "71 percent autonomous synthesis success" figure exists only in a sponsored article. That is not an accusation, it is a rung.

Mathematics: real results, and a scoreboard that flatters

The mathematics column is the strangest, because the results are unambiguous and the accounting is not. Gemini Deep Think's 35 of 42 at IMO 2025 was graded by IMO coordinators; OpenAI's identical score the same week was graded by former medallists the company hired, which is not the same thing. For IMO 2026 in July, the official IMO page mentions no AI entrants; a Hong Kong newspaper reports a Chinese lab's model scored 42 of 42 "graded by IMO", a GitHub repository holds Lean proofs of all six problems whose per-problem times sum to about 25 hours against a nine-hour human window, and several 42 of 42 claims for frontier models come from a harness graded by other models. I could verify none of the 2026 claims and I would not repeat them without that caveat.

On open problems the story is better and the bookkeeping is worse. On 20 May 2026 an OpenAI model produced a counterexample to Erdős's unit-distance conjecture, and a group including Alon, Bloom, Gowers and Wood published the verification, while attributing the underlying ideas to earlier human work. DeepMind's Aletheia swept 700 Erdős problems and reports four fully autonomous solutions. Terence Tao maintains a wiki of AI contributions whose recurring note is that many "solutions" are rediscoveries of results already in the literature; DeepMind itself counts nine of its contributions as "forgotten solutions found". And Thomas Bloom's complaint that AI-generated proofs run to 100 or 200 pages that "no human is going to read" is the one that will matter for the field. Tao's ICM essay from August is the balanced read. None of these models is public, so the mathematics column sits on the "no open weights" rung even as it produces the field's most verifiable results, because Lean does the verification.

Agentic science: one Nature paper, no accepted paper

The category my prospects ask about is the agent that does the science end to end. Google's Co-Scientist, a Gemini-based system that generates and ranks hypotheses, went from a February 2025 preprint with three wet-lab validations to a Nature paper announced at I/O in May 2026 with six external lab case studies. That is the strongest evidence in the column. FutureHouse's Kosmos reads roughly 1,500 papers and writes 42,000 lines of analysis code per run, and its own evaluation found 79.4 percent of report statements accurate, with seven discoveries of which four were novel and three reproduced unpublished work. A one-in-five error rate on statements is the number to remember when someone proposes to run this unsupervised.

The peer-review milestone has not happened. Sakana's AI Scientist-v2 got one of three papers into an ICLR 2025 workshop and withdrew it; Sakana said none met a main-track bar, and I found no main-track acceptance by any fully automated system as of this month. The Agents4Science venue in October 2025 accepted 48 of 253 AI-authored submissions, which is a conference designed to accept them. Meanwhile LABBench2, with 1,900 tasks, dropped frontier model scores by 26 to 46 percent relative to the first version, and a study of fabricated citations found 2,564 papers in 2025 with at least one fake reference, twelve times the 2023 rate, with 1.6 percent corrected. Both OpenAI (Prism, January) and Anthropic (Claude Science, June, with a hardware standard and $30,000 grants) now sell into this, which is where my sandbox question comes from: a science agent that writes and runs 42,000 lines of code per task needs a kernel boundary, and I have written enough about why.

The money and the machines

The demand side is government-scale now. The US Genesis Mission, ordered on 24 November 2025, reported over $5 billion in federal commitments and more than $800 million from 63 partners on 22 July 2026. Argonne's Solstice is 100,000 Blackwell GPUs with a 10,000-GPU Equinox due in the first half of this year; NERSC's Doudna had its pilot delivered in January. UKRI put £1.6 billion into AI for 2026 to 2030 with up to £137 million labelled AI for Science; the EU's RAISE programme offers €600 million of compute access. The private side I listed above: Isomorphic $2.1 billion, Lila $550 million, Periodic $300 million, CuspAI reportedly $450 million, Genesis Molecular's $120 million Incyte deal. Against that, the validated outputs are one Phase III trial, two operational weather systems, a handful of confirmed enzymes and phages, and a growing pile of Lean-checked mathematics. That gap is not a scandal. Drug development takes a decade and the models are three years old. It is just the number to hold in your head when reading the next release.

Running one of these

Boltz-2 is the one I would start with, because it is MIT-licensed, predicts a protein-ligand complex and a binding affinity in one call, and uses a public multiple-sequence-alignment server so you need nothing but a CUDA GPU. Version 2.2.1 is the current release.

pip install "boltz[cuda]==2.2.1"

cat > complex.yaml <<'EOF'
version: 1
sequences:
  - protein:
      id: A
      sequence: "<your protein sequence, one line, from UniProt>"
  - ligand:
      id: B
      smiles: "CC(=O)Oc1ccccc1C(=O)O"      # aspirin
properties:
  - affinity:
      binder: B
EOF

boltz predict complex.yaml --use_msa_server --out_dir out/
# out/boltz_results_complex/predictions/complex/ -> CIF model, confidence_*.json, affinity_*.json

If the licence on the weights matters to you, which it will if the output feeds a commercial pipeline, pip install openfold3==0.5.0 && setup_openfold gets the Apache-2.0 model instead, a year behind on training data. For weather, pip install anemoi-inference and ECMWF's published checkpoints. For genomes, pip install evo2. All of it fits on one GPU, which is the least-reported fact in this field: the science models are small, the training data was the expensive part, and the labs that built them gave much of it away.

My reading, for the prospects. The field has produced things that are real and checked: forecasts a national centre trusts, an enzyme that catalyses, a phage that replicates, a counterexample nine mathematicians signed. It has also produced two million crystals nobody has made, benchmark wins on benchmarks the winner wrote, and papers that rediscover the literature at 200 pages a time. The rung is the thing to ask about. If the answer is "our benchmark", wait a year.


Related: Open Weights in 2026, Robot Foundation Models in 2026 and Agent Observability in 2026.

I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.

Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.


Related

Sep 8, 2026

Robot Foundation Models in 2026: 500,000 Hours of Human Video, 50-Step Action Chunks, and a Winning Score of 0.26

What vision-language-action models actually do, from π0's flow-matching action head to Gemini Robotics 2 and GR00T N1.7; who shipped humanoids (Unitree 5,500, AgiBot 15,000) and who didn't (Tesla, zero); the teleoperation behind the demos; and the benchmark numbers that say a home robot is not close. With a LeRobot dataset I loaded this morning.

14 min
Sep 9, 2026

AI Power in 2026: 485 TWh, a 0.3 GW Stargate, a 116 GW Turbine Backlog, and 1,800 MW That Dropped Off the Grid in Seconds

The electricity numbers behind AI, kept honest: IEA's 485 TWh for 2025 and 950 by 2030, LBNL's 4.7 percent of US power heading to 12, what a prompt actually costs (0.24 Wh at Google, 30 times more in reasoning mode), which gigawatt campuses are energised versus announced (Stargate Abilene at 0.3 of 1.2 GW), where the power comes from (a 116 GW gas-turbine backlog with 2031 slots, nuclear restarts in 2027, one SMR construction permit), what the grid operators are doing about 233 GW queues and 1,800 MW load-loss events, rack power from 132 kW to 600 kW, and a script to measure your own job's energy.

14 min
Sep 9, 2026

On-Device Inference in 2026: 160 Tokens a Second on a Laptop, a 2-Bit Model in Every iPhone, and the Bandwidth Arithmetic That Decides What Stays Local

What actually runs on the device now: Apple's 3B model trained at 2 bits and the WWDC26 framework that lets any model plug in, Gemini Nano inside AICore, Windows ML and the quiet replacement of Phi Silica, the NPU TOPS numbers and why they are the wrong metric, the small-model field from Gemma 4 E2B to Qwen3.5-2B to LFM2.5 with licences, measured decode speeds from Google's own tables and an independent phone benchmark, the runtimes (llama.cpp with multi-token prediction, LiteRT-LM, ExecuTorch 1.0, MLX), the memory-bandwidth arithmetic that predicts all of it, and the model-extraction attacks nobody has fixed. With commands to run one today.

14 min