Dwarkesh Patel · Beren Millidge · 2026-09-11
Discussing what could prevent a superintelligence-transformed world by 2036, the participants describe the case for recursive self-improvement: an AI research agent only marginally better than humans, run in parallel across hundreds of thousands of instances on ever-faster chips, should eventually outweigh every other bottleneck. One guest calls recursive self-improvement a cumulative task, arguing a training program short enough to fit under a million tokens of Python could in principle train a model capable of it from scratch, since discoveries such as attention, mixture of experts and GRPO become fixed once made. The group counters that such a loop still needs an AI able to propose objectives and revise them without going off the rails for a long time, which current reinforcement learning, aimed at objectives already well specified, has not demonstrated. Chess-engine Elo, rising linearly since the 1980s before a discontinuity where engines passed the human range, is offered as an argument for a similar discontinuity in AI research capability. One guest calls models plausibly a millionfold behind humans in data seen from birth to adulthood against data seen during training.
The host repeatedly asks whether the mechanisms described match the scale of the jump in capability, asking why reinforcement learning, which changes only a small number of bits in a model, has been more successful than expected. One guest answers that most of the gain comes from mid-training on synthetic reasoning data, which takes a model most of the way to its final RL checkpoint before RL adjusts the policy through a few high-signal bits; another cites a benchmark called EdgeBench, said to show the task length a model can complete doubling every three months, as evidence of generalization in task horizon rather than reasoning. The host separately asks why Sonnet 5 and Opus 5 reportedly trail GLM-5.3 and Kimi K3 despite access to logit distillation; the guests answer with a difficulty-versus-realism trade-off in training environments and a possible uncanny valley in imitating a larger teacher. In the closing predictions, the host presses hardest when two guests put a tenfold productivity increase for AI researchers at five to ten years, longer than their own one-to-three-year estimate for a general remote worker, and one resolves this by saying the remote-worker estimate assumes competence rather than creative research ability.
Several concrete results are cited. A study run with a student named Jerry Han trained every open-source architecture recipe since 2019 against every dataset from the same period, finding that better data explains roughly a twelvefold compute-efficiency gain and architecture improvements roughly a 3.7-fold gain at small scale, well short of a roughly two-thousand-fold cumulative gain implied by an outside estimate of threefold-per-year progress since 2019, a gap one guest attributes to scale effects not visible at the sizes tested. The group cites an error in the original Kaplan scaling-law paper, which did not account for learning-rate annealing across checkpoints, as something that could have been caught years earlier, and cites muP, the scaling of learning rate with model width, as a similar result reachable by thought once the objective is fixed. One guest reports a language model called Talkie, trained only on data up to 1930, was fine-tuned on modern coding-agent data and outscored Claude 3 Opus on SWE-bench, while a separate paper reportedly found a model trained through fifth-grade mathematics could not be reinforcement-learned directly to college mathematics, though it could climb successive grade levels. Cursor's Composer model is described as running online reinforcement learning from users' accept-or-reject signal on completions, deploying an updated model every five hours if it improved on an internal benchmark called CursorBench.
Q “If we are in 2036 and we do not have billions of crazy superintelligences running around that have radically transformed the world, what is the most likely reason that does not end up being the case?” 0:20
“To have any kind of self-propelling automated loop, we need the AI to propose objectives, optimize them, figure that out, propose a new objective, and have this not go off the rails at any point for a long, long time.” 13:43
“They are plausibly a millionfold behind humans in terms of how much data a human sees from birth to adulthood versus how much a model sees from cold start to finishing training.” 46:45
“It is theoretically possible to have a less-than-a-million-token Python file which, from scratch, trains a model that is capable of recursive self-improvement.” 47:32
Q “Why has RL been more successful than one would have naively thought?” 1:18:22
Q “That is far away. You think it is longer than for a general remote worker?” 1:32:10