Model Collapse in 2026: The Loop Nobody Runs, a 40% Synthetic Phi-4, a 49.9% AI Web, and 300T Human Tokens Left
What the model-collapse research tested versus the headline: Shumailov's Nature loop deletes real data every generation and takes OPT-125m from 20 to 28 perplexity, while accumulating data bounds the error; my 100-generation Gaussian experiment shows median variance falling from 0.98 to 0.36 under replacement and holding at 0.98 under accumulation; how Phi-4 (40% synthetic), Llama 3, DeepSeek-R1 (800k samples), Qwen3, Kimi K2 and Nemotron-CC-v2 (38%) actually use it; Epoch's 300T-token stock, Graphite's 49.9% AI web, Cloudflare's 38,065:1 crawl ratio and the $1.5B Bartz settlement.
15 minOct 5, 2026MLOps in 2026: Evals Are the New Unit Tests, Judges Agree With Experts 66% of the Time, and OpenAI's Evals API Shuts Down on 30 November
The 2026 state of MLOps, checked against primary sources: OpenAI is closing its Evals API (read-only 31 October, off 30 November) and winding down fine-tuning while buying Promptfoo; Langfuse went to ClickHouse and W&B to CoreWeave; GDPval's model grader agrees with experts 66% of the time against 71% for humans; ARC-AGI-3 went from 0.51% to 99.9% in five months; Terminal-Bench retires a task when every frontier model solves it 5 of 5 times; MLflow 3.16, Kubeflow 26.03, Feast 0.66, lakeFS under BSL, and a pytest gate on a Wilson lower bound that I actually ran.
14 min