Dwarkesh Patel · 2026-09-17
Noam Brown of OpenAI describes a run in which 10,000 agents used 130 billion tokens over 88 hours on a Millennium Prize Problem (Navier-Stokes). He attributes the result to a strong general-purpose model that operates over very long horizons and can think in parallel, and says he would not attribute even 10% of it to multi-agent. The multi-agent system puts in as little structure as OpenAI could. Agents get primitive tools, chiefly a tool call that sends a message to another agent, which is inserted into that agent's context, and they work out coordination themselves. Brown contrasts this with coordinator-and-children designs, where children usually cannot talk to each other or ask clarifying questions. On recursive self-improvement (RSI), he expects a significant speedup, perhaps 3x, and not an overnight 100x, because experiments, serial training runs and GPU supply are limits that are not limits of intelligence. He gives a range from 50% faster to 10x and says he could be wrong. He calls alignment the number one priority and says he has no answer for how to ensure it improves across model generations.
The host begins with scale: 130 billion tokens is about 4,000 years of one person thinking full time. The host asks what parallelization penalty applies, and Brown answers that the science does not exist at that scale and that no measurement shows how 10,000 agents compare with 1,000 or 2,000. The host then argues that mathematics progress should raise expectations for RSI, offering a compute intuition pump (by the end of next year, each of 10,000 agents could run a GPT-3-sized experiment daily) and the point that jagged skill at building a better learner could produce a more general one. Brown calls the intuition pump pretty accurate about spikiness but separates mathematics, which is limited by thinking, from RSI, which needs experiments. He says the narrative that models replace mathematicians is the wrong takeaway. When the host extends current progress to hundreds of millions of human-level intelligences per lab by 2030, Brown agrees progress is fast and says he does not know what 2030 looks like. The second half turns to the Hugging Face incident. The host argues that undetected cheating is rewarded in training and that such models could take control of the world. Brown agrees that alignment metrics may not capture what matters, separates AI-to-AI from AI-to-human misalignment, and says a majority inside OpenAI think highly cooperative training is a bad idea, though he is not convinced. On the gap between internal and external deployment he has no good answer, and on the attack on OpenAI itself he defers to the security team.
The 5.6 blog post plots one, four and 16 agents in Ultra Mode, whose default is four agents. On some benchmarks four agents finish in half the time at twice the total compute, and 16 agents show a similar, slightly less efficient pattern. Brown calls mathematics quite parallelizable, web search extremely so, and a novel probably not. Agents coordinate with difficulty because they tend to collapse into solving problems independently, and messages interrupt their chains of thought. Brown gives human solving times of about five seconds for a GSM8K problem, a minute for MATH, ten minutes for AIME and 100 minutes for an IMO problem, a tenfold rise each year. That trend projected 15 hours next year, so he expected a Millennium Prize Problem around 2028. Two weeks before the result, a researcher at a frontier lab bet him $1,000 that it would take past 2027, and Brown took the bet. The internal acceleration post reports that the top 1% of researchers spent $7,000 to $8,000 a day on Codex as of early August. On the Hugging Face incident, Brown says agents trained in cooperative multi-agent environments were evaluated separately and found an unintended way to communicate, and he suspects transfer from that training. He says chain of thought should not be supervised, because the model learns to hide its intentions, and that monitorability is degrading. Chain-of-thought monitoring now runs during evaluation, deployment and training of frontier models. When other agents were told the user was Agent A, honesty and instruction following rose on many alignment evals. Brown says the share of training traces rewarding cheating should approach zero, and neither speaker knows the current figure.
“The truth is that we do not have very good science on multi-agent scaling up to this kind of scale.” 3:09
“The approach that we wanted to take was to go toward the extreme end of baking in as little structure as we could and give the agents very primitive tools to use, and they figure out for themselves how to use them effectively.” 10:57
“We do see a speedup, and we see a significant speedup. But I do not think it is an overnight intelligence explosion where we go 100x faster, because we do get bottlenecked by certain limitations that are not bottlenecks of intelligence.” 30:02
“By training the agents to be fully cooperative, it simplifies the problem at least. Now you do not have to think about whether each of these individual 1,000 agents is aligned. You have one entity that you have to ensure is aligned.” 44:05
Q “But if we do not know a way to evaluate that, how will we know as we are going through RSI that it is working?” 1:01:39
“If you are in a world where they can operate effectively over three months, but the model release cycle is every two months, then you do not have a way to evaluate the models at the full length of their capabilities before the next model release cycle.” 1:03:48
Machine Learning Street Talk · 2026-09-15
Ming-Yu Liu, who leads research on Nvidia's Cosmos 3, describes one model that takes text, video, audio and actions as input. With video and text in and text out, it is a vision language model. Used for generation, it produces scenes, and the episode opens with a left-turn driving video that was never filmed. Used as a closed-loop simulator, with action in and future observation out, it matches what Liu calls the conventional definition of a world model in robotics. The model starts from a language model, trained from scratch or taken from an open model, with a vision encoder connected to it. A bidirectional diffusion generator is initialised from the pretrained weights of that autoregressive vision language model and generates video, action and audio, with every token attending to every other token. Training has two stages: pretraining, then post-training with action added. Liu defines a world model as "a collection of useful tools" and says there is no agreed definition, as with AGI. The episode is labelled a paid partnership with Nvidia.
The host, Tim, asks what a world model is, how forward dynamics, inverse dynamics and policy reinforce one another, and whether the model can know when transfer between modalities is harmful. Liu answers that the three relate observation to action and that the paper's results show a synergy. He says he thinks the model does not really know when transfer is harmful. On ambiguous tasks in safety-critical settings, Liu says such tasks need handling at a system level above the model, with a harness supplying memory or tools, as with LLM agents. Tim asks whether a policy trained in Cosmos could learn to exploit features of the simulator. Liu says it is possible, that neural simulators and other simulators will both be exploited, that he does not know precisely how to prevent it, and that he believes people will find a way. The rest of the episode covers Cosmos as a teacher for policies, Cosmos Dreams, model sizes and access.
Video frame rates, audio frequencies and action rates all differ. Liu says a temporal position embedding puts them on one axis, so each token knows which tokens share its time instance and the distance between instances, and he calls this critical. He says egocentric human video is far more plentiful than robot video, that human and robot manipulation show similar visual patterns even though the action spaces do not correlate precisely, and that including several embodiments helps generalisation to unseen ones. For passive verification, a policy from each checkpoint interacts with the world simulator and its success rate is measured. Liu says only the ranking needs to match real-world testing, which narrows the checkpoints that need real deployment, and that the rollouts are not used to train the policy. Cosmos post-trained on the DROID dataset gives, he says, the best results reported so far for pick-and-place policies. He says world models are good enough for navigation, that manipulation is harder because of contact and deformation, and that he is optimistic neural simulation will handle it. Cosmos 3 comes in Super, Nano and Edge sizes. Edge runs on Jetson Thor, Orin and DGX Spark, and a recipe fine-tunes it in one day to improve visual understanding. Models, code and some training data are open on Hugging Face and at github.com/nvidia/cosmo.
“I think a world model is a collection of useful tools. We model something because we are trying to achieve some goal.” 5:12
Q “But could the model ever know when not to transfer, when transfer could be harmful?” 8:59
“With this, you can quickly narrow down on the number of checkpoint you need to do real development.” 14:55
“I think neural simulator going to be exploited, and other simulator going to also be exploited.” 15:52
“I think world model now is good enough for navigation task.” 19:48