Model Collapse in 2026: The Loop Nobody Runs, a 40% Synthetic Phi-4, a 49.9% AI Web, and 300T Human Tokens Left
A fair share of the text and code published on the internet in 2026 was produced inside sandboxes like mine. PandaStack, the Firecracker microVM cloud I run, exists so that agents can generate, execute and publish without supervision, and some of what they publish gets crawled and fed to the next model. So if recursive training really degrades models, I am part of the plumbing. I went back to the papers the headlines cite, ran the smallest version of the experiment myself, and read the labs' published recipes. The collapse result is real. The headline is wrong about what it applies to.
The loop the Nature paper actually tested
The paper everyone quotes is Shumailov, Shumaylov, Zhao, Gal, Papernot and Anderson, posted to arXiv in May 2023 as The Curse of Recursion and published in Nature in 2024 as "AI models collapse when trained on recursively generated data". The abstract says model-generated training content "causes irreversible defects in the resulting models, where tails of the original content distribution disappear". Early collapse is when "the model begins losing information about the tails of the distribution"; late collapse is when it "converges to a distribution that carries little resemblance to the original one, often with very small variance".
The experiment is specific. They fine-tune OPT-125m on wikitext2, generate a new training set from the fine-tuned model, and train the next generation on that. In the main scenario each generation sees the previous generation's output and "no original data", and perplexity degrades "from 20 to 28 perplexity points". In the second scenario a random 10 percent of the original data is resampled into every generation, with "only minor degradation of performance". The one-dimensional Gaussian analysis gives the mechanism: with a constant sample size M per generation, the estimate's variance about the truth grows as σ²(1 + n/M), a random walk that compounds.
The headline result is a property of a loop in which every generation discards all previous data, real data included. Nobody trains a production model that way, and the paper's own second scenario shows that keeping a tenth of the real data mostly fixes it.
Accumulate, don't replace
Gerstgrasser, Schaeffer and colleagues made this the central point in April 2024 with Is Model Collapse Inevitable?. When synthetic data replaces real data, performance degrades; when it accumulates alongside the real data, collapse is avoided across language models, diffusion models and VAEs, and their linear-model theory shows that under accumulation "test error has a finite upper bound independent of the number of iterations".
The strongest negative results are from Dohmatob, Feng, Kempe and co-authors. A Tale of Tails, ICML 2024, shows synthetic data truncates the tails and changes the scaling law itself, on arithmetic transformers and Llama 2. Strong Model Collapse, October 2024, argues that "even the smallest fraction of synthetic data (e.g., as little as 1% of the total training dataset) can still lead to model collapse", and that larger models can amplify it. Model Autophagy Disorder, from Baraniuk's group in 2023, put it as: "without enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decrease". Guo et al., The Curious Decline of Linguistic Diversity, NAACL 2024 Findings, built lexical, syntactic and semantic diversity metrics and found "a consistent decrease in the diversity of the model outputs through successive iterations" even where benchmark scores held.
These results agree more than the commentary suggests. Replace the real data and you collapse. Keep it and the error is bounded but not zero: synthetic tokens are worth less than human ones, and the first thing you lose is the tail and the diversity, which benchmarks do not measure.
Running the loop myself
I wanted both regimes side by side with numbers I had produced, so I wrote the smallest possible version: a Gaussian fitted by maximum likelihood to 200 samples, resampled and refitted for 100 generations, over 500 seeds. "Replace" fits only on the newest synthetic samples, the Shumailov loop. "Accumulate" fits on everything ever generated, real data included, the Gerstgrasser loop. The third mode adds a fresh 10 percent of real samples each generation, as in the Nature paper's second scenario. Variance is the diversity metric; the share of samples beyond two sigma is the tail.
# collapse.py: the smallest possible model-collapse experiment.
# A "model" is a Gaussian fitted by maximum likelihood to N samples. Each
# generation draws N fresh samples from the fitted model and refits.
# replace: fit only on the previous generation's samples (the Shumailov et al. loop)
# accumulate: fit on every sample ever produced, original real data included
# (the Gerstgrasser et al. loop)
# fresh10: fit on the new samples plus a fresh 10% of real samples each generation
# Variance is the diversity metric; the 2-sigma tail mass is what "losing the tails" means.
# Python 3.14.6, numpy 2.5.3. Deterministic per seed; summarised over SEEDS runs.
import numpy as np
N, GENS, SEEDS = 200, 100, 500 # samples per generation, generations, repeats
REPORT = (0, 10, 25, 50, 75, 100)
def fit(x): # MLE Gaussian (ddof=0), the paper's estimator
return x.mean(), x.std()
def run(mode, rng):
real = rng.normal(0.0, 1.0, N) # generation 0: real data from N(0, 1)
pool = real.copy()
mu, sd = fit(real)
var, tail = [sd**2], [np.mean(np.abs(real) > 2)]
for g in range(1, GENS + 1):
synth = rng.normal(mu, sd, N) # the model's own samples
if mode == "replace":
train = synth
elif mode == "accumulate":
pool = np.concatenate([pool, synth]); train = pool
else: # fresh10
train = np.concatenate([synth, rng.normal(0.0, 1.0, N // 10)])
mu, sd = fit(train)
var.append(sd**2); tail.append(np.mean(np.abs(synth) > 2))
return np.array(var), np.array(tail)
for mode in ("replace", "accumulate", "fresh10"):
V, T = zip(*(run(mode, np.random.default_rng(s)) for s in range(SEEDS)))
V, T = np.array(V), np.array(T)
print(f"\n{mode}: variance mean / median, 2-sigma tail mass median; N={N}, {SEEDS} seeds")
for g in REPORT:
print(f" gen {g:3d}: var {V[:, g].mean():.3f} / {np.median(V[:, g]):.3f} tail {np.median(T[:, g]):.4f}")
print(f" runs with variance below 0.5 at gen {GENS}: {(V[:, -1] < 0.5).mean():.1%}")
print(f" runs with variance below 0.1 at gen {GENS}: {(V[:, -1] < 0.1).mean():.1%}")
These are the real outputs.
| Generation | Replace, median variance | Replace, mean variance | Accumulate, median | 10% fresh real, median | Replace, median 2σ tail mass |
|---|---|---|---|---|---|
| 0 | 0.982 | 0.988 | 0.982 | 0.982 | 0.045 |
| 10 | 0.884 | 0.938 | 0.971 | 0.948 | 0.040 |
| 25 | 0.765 | 0.873 | 0.973 | 0.939 | 0.035 |
| 50 | 0.613 | 0.772 | 0.976 | 0.938 | 0.015 |
| 75 | 0.480 | 0.696 | 0.976 | 0.955 | 0.010 |
| 100 | 0.363 | 0.602 | 0.976 | 0.940 | 0.005 |
Under replacement, 62.6 percent of runs ended generation 100 with less than half the original variance and 10.4 percent with less than a tenth; the median tail mass beyond two sigma fell from 4.5 percent to 0.5 percent, which is "losing the tails" in one number. Under accumulation nothing happened: not one of 500 runs dropped below 0.5. The mean under replacement falls more slowly than the median because the random walk also produces a few runs where variance grows.
A Gaussian is not a language model, but the mechanism the theory papers prove and the Nature paper demonstrates on OPT-125m is this one: the catastrophe is a property of deleting your real data.
What 2025 and 2026 added
The follow-ups have moved from "does it happen" to "under what mix, with what filter, measured how". The most useful empirical result is Kang et al.'s Demystifying Synthetic Data in LLM Pre-training at EMNLP 2025, "a large-scale empirical investigation (>1000 LLMs with >100k GPU hours)". One third rephrased synthetic web text with two thirds natural text reaches the same validation loss 5 to 10 times faster at larger budgets; training only on generated textbook-style data "results in notably higher loss on many downstream domains especially at small data budgets"; the good ratio converges to about 30 percent rephrased synthetic. On collapse the evidence is "mixed": rephrased data showed none, pure textbook-style mixtures showed the pattern.
On filtering, Yi, Liu, Cheng and Xu's Escaping Model Collapse via Synthetic Data Verification, October 2025, revised July 2026, shows that verifier-guided retraining "can yield near-term improvements, but ultimately drives the parameter estimate to the verifier's 'knowledge center' in the long run", and "unless the verifier is perfectly reliable, these early gains will plateau and may even reverse". A verifier is a second model; its blind spots become the next generation's distribution.
Two 2026 papers measure what benchmarks miss. Proskurina, Gourru and Velcin's Fairness Collapse, August 2026, finds that under repeated synthetic training "fairness degradation emerges before substantial degradation is reflected by standard language-modeling metrics". Russell et al.'s How Much Is an AI Token Worth?, September 2026, is the first scaling-law paper for AI text found in the wild: using Pangram on FineWeb, "27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August"; for data-limited models the benefit of AI tokens "saturates as more are added and quickly reverses into harm", and for models with large human-text budgets "AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it".
There is also a retrieval version of the loop. Graphite's AI search collapse study, June 2026, ran 1,528 simulations over 1,019 prompts on three frontier models, feeding each model's own answers back as retrieved references: 79.6 percent ended in collapse to a single answer, and the models cited their own earlier outputs 38.9 percent of the time against 7.4 percent for human-written originals. That is the replace loop at inference time, in days rather than model generations.
How the labs actually use it
Against that background the published recipes are consistent: synthetic data is generated by a stronger model, filtered by a verifier or a judge, rephrased rather than invented, and mixed with real data that is never thrown away.
| Model or dataset | Synthetic data use | Numbers | Licence | Source |
|---|---|---|---|---|
| Phi-4 (Microsoft, Dec 2024) | GPT-4o generates ~50 types of synthetic data throughout pretraining | 14B params, ~10T tokens, ~400B synthetic; mix 40% synthetic, 15% web, 15% web rewrites, 20% code, 10% acquired; GPQA 56.1 vs GPT-4o 50.6, MATH 80.4 vs 74.6 | paper CC BY 4.0 | arXiv 2412.08905 |
| Phi-4-reasoning / -plus (Apr 2025) | SFT on o3-mini reasoning traces, then GRPO on verifiable maths | 1.4M prompt-response pairs, 8.3B tokens; RL on ~6K problems; AIME 2025 63.1 / 78.0 vs DeepSeek-R1 70.4, o3-mini-high 82.5 | arXiv 2504.21318 | |
| Llama 3 (Meta, Jul 2024) | Six rounds of rejection sampling with the 405B model; execution-feedback code synthesis | 15.6T pretraining tokens; ~2.7M synthetic code SFT examples; "most of our training data is model-generated" | Llama 3 licence | arXiv 2407.21783 |
| Llama 4 (Apr 2025) | Maverick codistilled from Behemoth; Llama judges prune easy SFT data | >30T tokens; >50% of SFT data removed as "easy" | Llama 4 licence | Meta blog |
| DeepSeek-R1 (Jan 2025) | RL with rule-based accuracy and format rewards, no neural reward model; distillation into six open models | ~800k SFT samples (600k reasoning, 200k other) into Qwen2.5 1.5B to 32B and Llama 3.1 8B / 3.3 70B; distilled 32B beat RL-trained 32B on AIME 2024, 72.6 vs 47.0 | MIT | arXiv 2501.12948, HF |
| Qwen3 (May 2025) | Qwen2.5-VL extracts PDF text; Qwen2.5-Math and -Coder synthesise textbooks, QA, code; strong-to-weak distillation | 36T tokens, 119 languages, "trillions" of synthetic tokens; small models at 1/10 of the GPU hours | Apache 2.0 | arXiv 2505.09388 |
| Kimi K2 (Moonshot, Jul 2025) | Rephrasing for token efficiency; agentic tool-use synthesis with simulated tools and a rubric judge | 1T / 32B active, 15.5T tokens; SimpleQA 23.76 raw at 10 epochs, 27.39 with 1 rephrasing, 28.94 with 10 rephrasings at 1 epoch; 3,000+ real MCP tools plus 20,000+ synthetic | arXiv 2507.20534 | |
| Nemotron-CC (NVIDIA, ACL 2025) | Rephrased Common Crawl plus classifier ensembles | 6.3T tokens; +5.6 MMLU over DCLM at 8B / 1T | arXiv 2412.02595 | |
| Nemotron-CC-v2 (Aug 2025) | Rephrasing, diverse QA, translated QA from Qwen3-30B-A3B and Mistral-NeMo-12B | 6,585.8B tokens, ~2,508B synthetic (~38%) | NVIDIA Data Agreement for Model Training | HF |
| Nemotron-Post-Training-Dataset-v1 (Jul 2025) | Reasoning traces from DeepSeek-R1-0528 and Qwen3-235B-A22B | 25,659,642 samples; 24.6M from R1-0528; STEM 20.7M, maths 2.0M, code 1.9M, chat 0.75M, tool calling 0.31M | CC BY 4.0 | HF |
| Nemotron-Personas-USA | Synthetic personas over US demographic fields (age, education, 567 occupations, 52 states and territories) | 1M rows | CC BY 4.0 | HF |
| Cosmopedia (Hugging Face, Mar 2024) | Mixtral-8x7B-Instruct textbooks and stories from web and curriculum seeds | 25B tokens, 30M files, >10,000 GPU hours; Cosmo-1B beats TinyLlama 1.1B, trails Phi-1.5 | HF blog | |
| SmolLM3 (Jul 2025) | Mid-training on open reasoning traces; preference pairs from Qwen3-32B vs Qwen3-0.6B | 3B params, 11.2T tokens; 35B reasoning tokens from OpenThoughts3-1.2M and Llama-Nemotron | HF blog |
Three things stand out. Nobody runs the replace loop: Phi-4's 40 percent synthetic sits on 60 percent web, code and acquired text, and Nemotron-CC-v2's 38 percent beside 62 percent filtered crawl. The synthetic data that works is either a rephrasing of real text (Nemotron-CC; Kimi K2's knowledge data, five SimpleQA points at a tenth of the epochs) or a verified trace in a domain with a checker (DeepSeek-R1, Phi-4-reasoning). And the open ecosystem is a distillation tree: DeepSeek-R1 traces fill 24.6 million of Nemotron's 25.7 million post-training samples, Qwen3 generates Nemotron-CC-v2 and SmolLM3's preferences, and the licences in the table, covered in Open Weights in 2026, are what make the tree legal. As far as I can find, neither OpenAI nor Anthropic publishes a data mix.
One caution on the verifiable-domain story: Shao et al.'s Spurious Rewards, June 2025, found Qwen2.5-Math-7B gains 21.4 points on MATH-500 from random rewards against 29.1 from ground truth, an effect that "often fail[s] to produce gains for other model families, such as Llama3 or OLMo2". Some of what looks like learning from verified traces is the optimiser surfacing what the base model already knew, and your eval has to tell the difference.
The supply side: how much human text is left, and who owns it
The human text is finite. Epoch AI's Will we run out of data?, June 2024, estimates the effective stock of public human-generated text at about 300 trillion tokens (100T to 1,000T) and projects training sets of that size "between 2026 and 2032", around 2028 under compute-optimal scaling and 2027 with 5x overtraining. Epoch's view of synthetic data there is sceptical: it "has only been shown to reliably improve capabilities in relatively narrow domains like math and coding". I found no updated estimate on Epoch's site in 2025 or 2026; Qwen3's 36T and Llama 4's 30T suggest the near end of that window is where the frontier already is.
The stock is also being diluted. Graphite's October 2025 study of 43,000 random Common Crawl URLs found AI-written articles overtook human-written ones in November 2024; its May 2026 update, 55,400 URLs and three detectors (Pangram, Copyleaks, GPTZero), put the AI share at 50.9 percent in Q4 2025 and 49.9 percent in Q1 2026, with false-positive rates of 1.36 to 1.84 percent. Ahrefs' April 2025 sample of 900,000 new pages found 74.2 percent contained some AI text but only 2.5 percent were purely AI; Originality.ai's series on top-20 Google results puts detected AI content at 17.31 percent in September 2025. Russell et al.'s 27.5 to 31.1 percent of FineWeb tokens is the number that matters for pretraining.
Then the owners started charging. Cloudflare's 1 July 2025 announcement changed the default for new domains to block AI crawlers unless they pay, and pay per crawl implements it with HTTP 402, a publisher-set per-request price and Ed25519-signed Web Bot Auth, still in private beta as of July 2026. Cloudflare's August 2025 data gives the crawl-to-refer ratios for July 2025: Anthropic 38,065:1, OpenAI 1,091:1, Perplexity 194:1, Google 5.4:1, with training at 79 percent of AI bot activity. The 2025 Radar review adds that AI bots were 4.2 percent of HTML requests, and that GPTBot, ClaudeBot and CCBot were the most fully disallowed user agents in robots.txt. CCBot being blocked should worry open-model builders: Nemotron-CC and FineWeb are downstream of it.
The courts set a price too. In Bartz v. Anthropic, Judge Alsup's June 2025 summary judgment held training on lawfully bought books to be fair use while the roughly 7 million pirated library copies were not, and the $1.5 billion settlement, about $3,000 per work across roughly 482,460 titles, got preliminary approval in September 2025 and final approval in July 2026. Those details are from the case summary rather than the docket, which I could not fetch; the New York Times case against OpenAI I could not verify from a primary page, so I leave its status out. At $3,000 a book, half a million books is $1.5 billion, which makes 400 billion synthetic tokens from your own model look cheap, and explains why every lab in the table is generating rather than buying.
Tabular synthetic data is a different product with a different problem
The other synthetic-data market, test and analytics data for regulated databases, is about privacy rather than collapse, and it consolidated this year. Gretel's domain now redirects to NVIDIA's synthetic data page, which lists NeMo Data Designer and a NeMo Safe Synthesizer that produces "privacy-safe versions of sensitive data" for HIPAA and GDPR; no price was disclosed and I found no primary announcement. MOSTLY AI ships its TabularARGN SDK under Apache 2.0 with "built-in differential privacy". Tonic and Synthesized sell masking, subsetting and generation for test environments; neither site makes a differential-privacy claim.
The research is less kind than the marketing. Stadler, Oprisanu and Troncoso's Synthetic Data: Anonymisation Groundhog Day, USENIX Security 2022, found that "synthetic data either does not prevent inference attacks or does not retain data utility", and that without a formal guarantee the privacy outcome is unpredictable. Differential privacy is the only property that survives that paper, and a vendor who says "DP" without stating epsilon has told you nothing. For a platform team, synthetic test data is a build artifact: pinned pipeline, epsilon and seed recorded, versioned and promoted through environments like anything else in CI.
When it helps, when it hurts, and what to measure
The rules are short. Synthetic data helps when there is a checker: maths with a final answer, code with tests, tool calls against a simulator, which is where DeepSeek-R1, Phi-4-reasoning and Kimi K2 all sit. It helps as rephrasing of real text at around a third of the mix. It hurts when it replaces real data, when the generator's or verifier's blind spots define the distribution, and in the tails: creative tasks (Guo et al.), demographic fairness (Proskurina et al.), and whatever your benchmark does not cover. KORMo, a 10.8B model pretrained on 68.74 percent synthetic Korean data without reported instability, says a low-resource language is not automatically a tail if the synthetic data is curated for coverage, but that curation is the work.
The metrics are the part most teams skip. Perplexity on held-out human text is the Nature paper's instrument and catches late collapse; it did not catch the diversity loss in Guo et al. or the bias drift in Proskurina et al. So measure lexical, syntactic and semantic diversity, precision and recall against a real reference set in Alemohammad's sense, per-group performance on the slices you care about, and the share of AI-generated tokens in your corpus with a detector whose false-positive rate you know, the way I argued for agent traces in Agent Observability in 2026. That last one is a platform problem: every synthetic token in your lake should carry its generator, version, filter and seed, and the real data should never be deleted to make room, because the real data is the asset.
What I take from this is that the headline inverts the finding. Models collapse when you throw away the real data, and the labs, who have read the same papers, accumulate, filter and rephrase instead. The real constraint is upstream: the human text is finite, half the new web is machine-written, the owners are charging by the request or by the lawsuit, and the value of an AI token, measured in the wild for the first time this September, turns negative sooner than anyone selling synthetic data will admit. For my own platform the step is small and I would do it this quarter: tag what agents produce on their way out so it carries its provenance, and keep the human data I hold under a retention policy that says never. Everything else is a mix ratio, and the right one is still an experiment.
Related: Evals Are the New Unit Tests, Open Weights in 2026 and LLM Inference in 2026.
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
Open Weights in 2026: Chinese Labs Ship the Frontier, the Licenses Grew Clauses, and Brussels Started Enforcing
DeepSeek V4, Kimi K3, GLM-5.3, Qwen3.8 and Xiaomi's MiMo are the open frontier now; Meta went closed and came back; the licences quietly acquired revenue thresholds; the EU AI Office got its enforcement powers on 2 August. What I found when I read every licence file, with the Hugging Face API calls that check them.
16 minOct 5, 2026FinOps for AI in 2026: 98% Manage Token Spend, 3 in 4 Can't Prove the Value, GPUs Run at 5%, and FOCUS Gets a Token Column in 1.5
The AI cost ledger as of October 2026: State of FinOps 2026 (1,192 practitioners, 98 percent manage AI spend, up from 31 percent in 2024), the Tokenomics Foundation finding that three in four enterprises cannot prove AI outcomes to the CFO, a per-million-token price table for Anthropic, OpenAI, Google and DeepSeek with cache and batch multipliers, a 20-step agent loop at $1.58 uncached and $0.42 cached, the FOCUS 1.2 to 1.5 roadmap and who exports which version, LiteLLM and Cloudflare budgets, Cast AI's 5 percent GPU utilisation, and $0.0037 of model cost inside a $0.99 resolution.
13 minOct 5, 2026GPU Clouds in 2026: H100s at $2.59 an Hour, $104B of CoreWeave Backlog on $35B of Debt, and a Six-Year Bet on Silicon
The GPU rental economy from filings and price pages fetched on 5 October 2026: H100 hours that fell from about $8 in 2023 to a $1.70 trough in October 2025 and recovered to $2.59 on the transacted index while list prices rose 25 percent; CoreWeave's $2.58B quarter, $103.7B backlog, $35.1B of debt from 15 percent DDTLs to a 5.9 percent A3 facility, and Microsoft at 67 percent of 2025 revenue; Nebius, Lambda, Crusoe, Oracle's $664B RPO and NVIDIA's $105B of guarantees; who books servers over 5, 5.5 or 6 years and what Burry's $176B means; and an own-versus-rent break-even calculator.
13 min