World Models in 2026: Genie 3 at 24 fps, Cosmos 3 at 64B, V-JEPA 2 Planning in 16 Seconds, and an $8.2B Exit to AMD

Oct 5, 2026 · 14 min · Ajay Kumar

My platform runs sandboxes for AI agents: a microVM boots in 179 milliseconds, the agent runs code in it, and the VM is thrown away. That is a world of a literal kind, where the rules are a Linux kernel. The other kind of world an agent can act in is a learned one, a network that takes a frame and an action and emits the next frame, and in 2026 that kind has become a product category with a billion-dollar seed round, an $8.2 billion acquisition, a $249.99 a month consumer tier and three open-weight releases from NVIDIA alone. I spent the week on the papers, model cards and licences, and ran the one model that fits on a laptop. This is what the term means now, what shipped, what it costs to run, and which uses pay.

Three things called a world model

The phrase has a precise 2018 origin and a loose 2026 usage. David Ha and Jürgen Schmidhuber's World Models compressed frames with a VAE, predicted the next compressed state with a recurrent network, and trained a tiny controller entirely inside the model's own "dream" before transferring it back to CarRacing and VizDoom. Everything since scales that loop; the lineages disagree on what the model predicts.

The first predicts pixels. GameNGen ran DOOM at over 20 frames a second on a single TPU in 2024 with human raters barely better than chance at spotting the simulation, and DeepMind's Genie line, from the 11-billion-parameter Genie of February 2024 onward, learned a latent action space from unlabelled video so a user could steer with keys the model had never seen labelled. The second predicts representations. Yann LeCun's 2022 position paper defines the world-model module as the part that estimates missing state and predicts plausible futures in an abstract space where "irrelevant details are ignored". That is JEPA. The third predicts explicit state a program can load: meshes, Gaussian splats, depth, sensor returns.

World Labs published the most useful framing in June, a functional taxonomy of renderers (pixels for human eyes, judged on fidelity; Genie 3 and its own RTFM), simulators (geometrically and physically faithful state; Marble) and planners (models that output actions, the least validated category). I will use that split because it maps onto who pays.

Three things called a world model, and what each one is for (2026) Renderer: pixels for human eyes Predicts the next frame from the last frames plus an action or a prompt. Genie 3: 720p, 20-24 fps, minutes RTFM: one H100, interactive fps Odyssey-1: 30 fps, 40 ms, $1-2/hour GWM-1: 720p, real time, 2 min HY-World 1.5: 480p, 24 fps, open Metric: looks right, stays consistent Pays: games, agent training worlds Fails: geometry drifts, no collider Simulator: state a program loads Outputs geometry, physics, sensors that engines and robots consume. Marble: splats, meshes, World API Atlas: 1440p video, splats, depth HY-World 2.0: 3DGS, meshes, Isaac Cosmos Transfer 2.5: sim to real GAIA-2: multi-camera driving video Metric: geometric, physical fidelity Pays: robot sim, AV testing, 3D Exit: AMD paid $8.2B for World Labs Planner: acts in latent space Predicts the next representation, not the next pixel; plans over it. V-JEPA 2-AC: 16 s per action vs Cosmos 4 min; 80% pick-and-place Cosmos 3: 64B/16B/4B, actions out AMI Labs: $1.03B at $3.5B, 5 years LeJEPA: 79% ImageNet, 50 lines Metric: task success, planning cost Pays: robot policies, agent planning Fails: hard to inspect, no pictures Counter-evidence:Physics-IQ: best model 29.5 of 100; realism vs physics correlation -0.46 (not significant) Veo 3: 78% of 5x5 mazes at pass@10 (Veo 2: 14%). WorldModelBench: 14 models, 67,000 human labels
Categories follow World Labs' June 2026 taxonomy; figures come from the model cards, papers and announcements linked in the text and listed in the table below.

What shipped, with dates and licences

Model Date Output Resolution / rate Access Licence
Genie 3 (DeepMind) 5 Aug 2025 interactive video 720p, 20 to 24 fps, a few minutes Project Genie, AI Ultra only, 18+, 140+ countries proprietary
RTFM (World Labs) 16 Oct 2025 interactive video interactive rate on one H100 demo proprietary
Marble (World Labs) 12 Nov 2025 Gaussian splats, meshes, video n/a web, World API from 21 Jan 2026 proprietary
Atlas (World Labs) 1 Sep 2026 images, video, splats, point clouds, depth up to 1 min at 1440p early access proprietary
Cosmos Predict 2.5 (NVIDIA) 6 Oct 2025 video from text, image or video 1280x704, 16 fps, 5 s Hugging Face, 2B and 14B Apache 2.0 code, NVIDIA Open Model License weights
Cosmos 3 (NVIDIA) 31 May 2026 text, image, video, audio, action sequences up to 720p, 5 to 400 frames Hugging Face, 64B, 16B, 4B OpenMDW 1.1
V-JEPA 2 (Meta) 11 Jun 2025 latent embeddings and predictions 256 or 384 px, 64 frames Hugging Face, 300M to 1B MIT
GWM-1 (Runway) 11 Dec 2025 interactive video 720p, real time, up to 2 min by request proprietary
HY-World 1.5 WorldPlay (Tencent) 17 Dec 2025 interactive video 480p, 24 fps GitHub and Hugging Face, 8B and 5B Tencent licence file
HY-World 2.0 (Tencent) 16 Apr 2026 3DGS, meshes, point clouds, depth n/a GitHub, staged releases Tencent licence file
Wan 2.2 TI2V-5B (Alibaba) 28 Jul 2025 video 720p, 24 fps, 5 s Hugging Face Apache 2.0
Odyssey-1 28 May 2025 interactive video 30 fps, 40 ms latency, 5+ min research preview proprietary

Two cautions. "Interactive video" hides the gap between rendering a prompt you typed and rendering a sensor stream a robot can train on. And the Tencent licence cells point at a file I could not read in full; if you ship on those weights, read the text, as the open-weights post argues.

Renderers: Genie 3 and the real-time crowd

Genie 3 is the reference point. DeepMind's announcement on 5 August 2025 gave 720p at 24 frames a second, "a few minutes" of consistent interaction with visual memory reaching back about a minute, and "promptable world events", text that changes the weather or drops an object into a running world. Its listed limits still stand: a constrained action space, trouble with multiple agents, no geographic accuracy, sessions in minutes rather than hours. Genie 2 had managed up to a minute with most examples at 10 to 20 seconds. Public access is Project Genie in Google Labs: the model page quotes 20 to 24 fps and Street View grounding, the Labs page restricts it to Google AI Ultra subscribers aged 18 and over, Google's plan comparison says 140-plus countries, and Ultra launched at $249.99 a month. I found no primary page giving the launch date, and nothing on a "Genie 4". The agent side is SIMA 2, announced 13 November 2025 with Gemini as its reasoning core; dropped into unseen Genie 3 worlds, it oriented itself and followed instructions, the 2018 loop at scale, still a limited research preview.

The rest of the real-time field is defined by one problem, stated plainly in the RTFM post: 4K at 60 fps is over 100,000 tokens a second, and an hour of persistent world means attending over more than 100 million tokens. RTFM keeps a spatial memory of posed frames, retrieves only those near the camera, and runs on one H100. Odyssey's first preview in May 2025 gave the only published cost I found, $1 to $2 per user-hour on H100 clusters at up to 30 fps and 40 ms latency, self-described as "raw, unstable"; its site now lists an Odyssey-3 and a $310 million Series B I could not verify elsewhere. Decart's Oasis did Minecraft at 20 fps in October 2024 with 500 million open parameters. Runway's GWM-1 of 11 December 2025 holds 720p in real time for up to two minutes and ships as Worlds, Robotics and Avatars by request. Tencent's WorldPlay is the open one: 480p at 24 fps, an 8B or 5B backbone, 28 to 72 GB of GPU memory.

Simulators: the ones that produced an exit

The renderer gives you a picture; the simulator gives you something a program can load, and that is where the money went. World Labs raised a $230 million seed in September 2024, shipped Marble on 12 November 2025 with export to Gaussian splats, collider meshes and camera-controlled video, opened the World API on 21 January 2026, and raised $1 billion on 18 February from a list including AMD, Autodesk, Fidelity, NVIDIA and Sea. On 1 September it announced Atlas, an "omni model" pretrained from scratch on text, images, video and 3D that emits up to a minute of 1440p video plus splats, point clouds and depth, preferred over FLUX in 93 percent of camera-control comparisons. Four weeks later AMD agreed to acquire the company for about $8.2 billion in stock, Fei-Fei Li becoming executive vice president and chief scientist, closing by year end. Marble's consumer pricing is on no page I could fetch.

The robotics case is World Labs' July real-to-sim-to-real post: reconstruct one physical task as an interactive world, vary appearance, clutter, physics and viewpoint into thousands of variants, train the policy entirely in simulation, evaluate on 2,000 simulated trials per checkpoint, deploy. Policies ran autonomously for an hour on cable routing and object transfer across five robot platforms, with 100 real trials per task. Tencent's HY-World 2.0, released in stages from 16 April, makes the same pitch for its 3DGS and mesh outputs, importable into Unity, Unreal and Isaac, and its August WorldClaw paper uses planning agents to turn a prompt into terrain, assets and editable meshes.

NVIDIA straddles the simulator and planner columns. Cosmos launched on 6 January 2025 with 1X, Agility, Figure, Waabi, Wayve and Uber as named users. Cosmos Transfer 2.5, 2B parameters from October 2025, is the sim-to-real piece: a multi-ControlNet that takes depth, segmentation, edges or lidar from an Omniverse render and produces photoreal video with the same geometry, which is how a synthetic scene becomes camera-like training data. Wayve's GAIA-2 is the driving equivalent, a latent diffusion model conditioned on ego speed and steering, agent boxes, weather and road layout across the UK, US and Germany, for the near-collisions no fleet should collect on purpose. On Waymo and Tesla I found no primary page describing a named world model, so both stay out. The robot side continues in the robot foundation models post.

Planners: JEPA, Cosmos 3 and a billion-dollar bet

The planner column is where the definitional fight is sharpest. Meta's V-JEPA 2, released 11 June 2025 under MIT, is a 1.2-billion-parameter encoder-predictor trained on over a million hours of video, then post-trained with less than 62 hours of unlabelled robot video from Droid into V-JEPA 2-AC. Deployed zero-shot on Franka arms in two labs, it reached 65 to 80 percent on pick-and-place of unseen objects, and the paper's comparison is the number for the slide: planning one action over 800 latent candidates took 16 seconds, while the pixel-generating Cosmos baseline took 4 minutes for 80 samples and scored 0 percent on grasp and pick-and-place in the same lab. If you only need to plan, rendering pixels is wasted compute; that is the whole JEPA argument. Meta followed with V-JEPA 2.1 on 16 March 2026, a 2B ViT-G at 384 pixels, and the LeJEPA paper by Balestriero and LeCun replaced the stop-gradients, teacher-student copies and schedulers with a 50-line objective reaching 79 percent ImageNet linear accuracy on a frozen ViT-H/14.

LeCun then left. He announced his departure from Meta on 19 November 2025, founded Advanced Machine Intelligence Labs in Paris on 15 December, and in March 2026 it raised $1.03 billion at a $3.5 billion pre-money valuation, Bezos Expeditions among the co-leads and NVIDIA and Samsung among the investors, with Alexandre LeBrun as CEO and a stated horizon of about five years. A billion dollars on an architecture whose strongest published result is 80 percent pick-and-place is a bet, not a business.

NVIDIA's answer is to stop choosing. Cosmos 3, released 31 May 2026 under the Linux Foundation's OpenMDW 1.1, is a Mixture-of-Transformers pairing an autoregressive transformer for discrete tokens with a diffusion transformer for continuous ones, so one checkpoint reasons over video, generates it and emits action sequences. It ships as Super at 64B for H200, B200 and GB200, Nano at 16B for an RTX Pro 6000 or H100, and Edge at 4B for Jetson Thor; Nano takes text, images, five frames of video, half a second of audio and 16 to 400 frames of action trajectory, trained on 1.3 billion items from 393 datasets. NVIDIA's August post names Doosan, LG, Samsung and Skild AI on robots and Li Auto and Xiaomi on vehicles. Whether one model doing all three beats three specialised ones is the question the taxonomy post was written to ask.

Running one on a laptop

Most of the table needs an H100. V-JEPA 2 does not. The script below is what I ran on an Apple M3 Pro with 18 GB, CPU only, torch 2.14.1, transformers 5.18.0, torchvision 0.29.1, pillow 12.3.0; the clip is synthetic so there is no video decoder to install.

# vjepa2_embed.py: pull V-JEPA 2 (ViT-L/16, MIT licence) from Hugging Face and get the
# encoder embedding plus the predictor's output for a 64-frame clip. CPU is fine.
# Pinned: torch==2.14.1 transformers==5.18.0 torchvision==0.29.1 pillow==12.3.0
import time, torch, numpy as np
from transformers import AutoVideoProcessor, AutoModel

repo = "facebook/vjepa2-vitl-fpc64-256"          # fpc64 = 64 frames per clip, 256px crop
processor = AutoVideoProcessor.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, attn_implementation="sdpa").eval()
n_params = sum(p.numel() for p in model.parameters())

T, H, W = 64, 256, 256
video = np.zeros((T, H, W, 3), dtype=np.uint8)   # T x H x W x C, uint8
for t in range(T):                                # a 48px square drifting 2px per frame
    x = 20 + 2 * t
    video[t, 100:148, x:x + 48, :] = 220

inputs = processor(list(video), return_tensors="pt")
t0 = time.time()
with torch.no_grad():
    out = model(**inputs)
dt = time.time() - t0

enc = out.last_hidden_state                       # [batch, tokens, 1024]
pred = out.predictor_output.last_hidden_state     # predictor output, same space
print(f"params: {n_params/1e6:.0f}M  input: {tuple(inputs['pixel_values_videos'].shape)}")
print(f"encoder tokens: {tuple(enc.shape)}  predictor: {tuple(pred.shape)}  forward: {dt:.1f}s CPU")
clip_vec = enc.mean(dim=1)                        # one 1024-d vector for the whole clip
print(f"clip embedding: {tuple(clip_vec.shape)}  norm {clip_vec.norm():.2f}")

Real output; the one-time weight download pushed the whole run to 1 minute 38 seconds of wall clock:

params: 326M  input: (1, 64, 3, 256, 256)
encoder tokens: (1, 8192, 1024)  predictor: (1, 8192, 1024)  forward: 26.0s CPU
clip embedding: (1, 1024)  norm 74.02

The 8,192 is 64 frames in tubelets of two times a 16 by 16 patch grid, so every token is a two-frame, 16-pixel-square patch of the world, and the predictor returns the same shape because the default masks ask it to predict every patch. Twenty-six seconds on a laptop CPU is one step of a planning loop that samples hundreds of candidates; the paper's 16 seconds per action was on a GPU with 800 of them. A GPU-less pip install is a lower bar than any renderer in the table.

The rendering costs are blunter. Cosmos Predict 2.5-2B needs 32.54 GB of GPU memory for 720p and about 229 seconds on an H100, or 124 on a B200, for a 5-second clip at 16 fps, roughly 46 H100-seconds per second of video: fine for an offline data pipeline, useless for interaction. Cosmos3-Super wants eight H200s for a 55-second generation at 50 steps, or two H200s at about three minutes. The real-time models never publish per-frame costs, which is itself information; the one number that exists is Odyssey's $1 to $2 per user-hour. NVIDIA's offline command, from its inference docs, is python examples/inference.py -i assets/base/robot_pouring.json -o outputs/base_video2world --inference-type=video2world, with torchrun --nproc_per_node=8 for the 14B.

Video models as world models, and the counter-evidence

The strongest version of the implicit-world-model claim is DeepMind's September 2025 paper, Video models are zero-shot learners and reasoners. Across 18,384 generated videos over 62 qualitative and 7 quantitative tasks, Veo 3 segmented objects at 0.74 mIoU, detected edges at 0.77 against a task-specific 0.90, extracted objects at 93 percent pass@10, and solved 78 percent of 5 by 5 mazes at pass@10 where Veo 2 managed 14. The thesis is that video models are on the LLM path, and Google's Veo 3.1 page sells "real world physics" as a feature. OpenAI has made the claim since its February 2024 report Video generation models as world simulators, and Sora 2 shipped on 30 September 2025 with that framing intact and the earlier admission of "limited capacity to simulate complex physics" still on record.

The counter-evidence is specific. Physics-IQ, also from DeepMind, filmed 66 physical scenarios from three angles twice each, 396 real videos, and asked models to continue them; the best score was VideoPoet's 29.5 against a ceiling of 100 set by the variance between two real recordings, Runway Gen 3 got 22.8, Sora 10.0, and the correlation between an MLLM's realism judgement and the physics score was negative 0.46 and not significant. WorldModelBench judged 14 video models with 67,000 human labels for instruction following and physics adherence, catching objects that change size mid-clip, and found that training against those labels improved the scores, so the failures are partly a reward problem rather than a ceiling. Meta's IntPhys 2 has humans at 85 to 95 percent and models near chance on telling plausible clips from impossible ones. "Zero-shot reasoner" and "does not understand physics" are both true of the same model: a maze is a visual pattern, a bouncing ball is a conservation law.

What pays, and the discipline it needs

Strip out the valuations and four uses have customers. Robot training data, where Cosmos Transfer, Marble's real-to-sim-to-real engine and GWM Robotics all sell more varied trajectories than a fleet can collect, with real evals (100 trials) still far below simulated ones (2,000). Driving simulation, where GAIA-2 and Cosmos's seven-camera variant generate the near-misses you cannot ethically record. Games and interactive media, which is what Project Genie, Marble, Odyssey, Decart and GWM Worlds are, at $249.99 a month or by request. And agent planning, V-JEPA 2-AC, SIMA 2 in Genie 3 worlds, Cosmos 3's action output: the fewest paying customers and the most research money.

All four are pipelines, which is where my own discipline applies. A world model used for training data is a dependency with a version, a licence, a dataset manifest and an eval that must rerun when the weights change. The scores that matter are the held-out physical trials, not the demo reels, and tracking a Physics-IQ or a real-robot success rate across versions is the same evals-as-tests argument I have made for agents. The GPU budgets are unambiguous too, 32 GB and four minutes a clip offline, a dedicated H100 per user for anything interactive, numbers for a capacity plan before they reach a roadmap.

What I take from the year. The word has split into three things and the money followed the one that outputs state a program can load: World Labs sold for $8.2 billion because Marble exports a mesh, not because RTFM looks good. The pixel renderers are real products with real limits, minutes of consistency, one GPU per player, no published unit economics. The latent planners have the best argument, the thinnest evidence and a billion dollars of patience. This quarter I would keep the V-JEPA 2 encoder in the toolbox as a cheap, MIT-licensed video embedding that runs on a CPU, the one piece of this field a small team can own; watch Cosmos 3 Edge at 4B for on-device inference; and refuse to put a world model in any pipeline without a held-out physical eval and a pinned version, because a simulator that is 29.5 percent right about physics and 93 percent convincing is the most dangerous kind of dependency there is.


Related: Robot Foundation Models in 2026, Open Weights in 2026 and On-Device Inference in 2026.

I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.

Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.


Related

Sep 9, 2026

AI for Science in 2026: One Phase III Trial, Two Operational Weather Models, and Erdős Proofs Nobody Wants to Read

What has actually been validated outside the lab that made the claim: rentosertib's Phase III start, ECMWF and NOAA running AI forecasts operationally, Evo 2's viable phage genomes, the Erdős-problem scoreboard and its rediscovery problem, the GNoME and MatterGen materials that turned out not to be new, and the licence and pip-install status of every open model an engineer can run this week.

16 min
Sep 9, 2026

On-Device Inference in 2026: 160 Tokens a Second on a Laptop, a 2-Bit Model in Every iPhone, and the Bandwidth Arithmetic That Decides What Stays Local

What actually runs on the device now: Apple's 3B model trained at 2 bits and the WWDC26 framework that lets any model plug in, Gemini Nano inside AICore, Windows ML and the quiet replacement of Phi Silica, the NPU TOPS numbers and why they are the wrong metric, the small-model field from Gemma 4 E2B to Qwen3.5-2B to LFM2.5 with licences, measured decode speeds from Google's own tables and an independent phone benchmark, the runtimes (llama.cpp with multi-token prediction, LiteRT-LM, ExecuTorch 1.0, MLX), the memory-bandwidth arithmetic that predicts all of it, and the model-extraction attacks nobody has fixed. With commands to run one today.

14 min
Sep 8, 2026

Agent Observability in 2026: The Spans Are Standard, the Standard Isn't Stable, and Nobody Reads the Logs

OpenTelemetry's GenAI conventions moved to their own repo in June and still carry the Development badge; every vendor from Datadog to Grafana ingests them anyway. I instrumented an agent loop with the official Python library, show the exact spans it emits, and go through what the 2026 eval papers say actually catches failures: pass^k, log inspection, and a step budget.

13 min