CLANKERS WEEKLY

Week of 2026-08-31 · 44 items

Unitree's Shanghai IPO surged 629% to a ~$66B valuation as SoftBank moved to take a majority stake in 1X at ~$6B and XPeng's Dogotix raised $900M+ for its IRON humanoid, while Bedrock Robotics ran operator-free excavators on live Texas and Nevada sites. Humanoid capital keeps compounding, but Amodei's warning that robotics is moving far slower than software AI is the check on that enthusiasm.

What happened 20

DeepMind's Shane Gu: video models already show zero-shot physical intuition robots need

AI Engineer · 2026-08-30

A frontier lab researcher argues video foundation models are becoming the missing world-model layer for embodied AI, not just content generators.

“It can do robotics — it has really good physical intuitions, like a world model. The key is really the mix of visual reasoning and text reasoning tied together.” 10:15

“Eight months ago we put out the evaluation paper called 'Video Models are Zero-Shot Learners and Reasoners.'” 9:48

“It's a missing foundational model that's absolutely required if you want to make an AGI that matches humans, not just a 'jacked' one.” 19:18

“We've announced publicly that we have some robotics collaboration — we have a robotics team at GDM that's always interested in things like this.” 47:35

“Right now the recipe that works is pre-training that scales a lot, and that's what learns a lot of intelligence.” 15:55

Tileubay: cloud robotics fails control loops — 50ms latency risks accidents vs 500ms fine for a chatbot

The Robot Report · 2026-08-30

Explains why offloading robot control to the cloud is a physics/latency problem, not just a bandwidth one.

“For a cloud-based text chatbot, a 500-millisecond delay goes unnoticed by the user. For a bipedal humanoid robot or an autonomous vehicle at an intersection, a latency of even 50 milliseconds carries a high risk of an accident.”

“Shifting the critical decision-making loop to the cloud means that even a minor packet loss or a temporary connection drop instantly turns the robot into an unguided physical object weighing dozens or hundreds of kilograms, posing an immediate threat to its surroundings.”

“Even under the extremely conservative assumption that a robot faces only 10 alternative options at each step, the size of the search space expands exponentially as planning depth increases.”

“In a series of controlled simulations and computational experiments, the CCE algorithm demonstrated an ability to compress the search space by a factor of 8 to 11 while preserving the functional quality of decisions, establishing a pathway toward enhancing edge AI efficiency.”

Tencent's Hy4 scales to 770B params (49B active) and 1M context, up sharply from Hy3

Simon Willison · 2026-08-29

Marks continued rapid parameter and context scaling among Chinese open-weight labs, relevant to compute demand.

“New open-weight, text-only LLM from Tencent: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face.”

“There are just two reasoning effort levels: "high" (the default) and "no_think" (reasoning disabled).”

“The reasoning trace uses slightly truncated English, presumably because perfect grammar isn't useful or token-efficient for hidden reasoning text.”

AWS's Strands lets an LLM agent choose which robot policy to run, not just execute one

AI Engineer · 2026-08-29

Shows an agentic orchestration layer over VLA policies as a stopgap before robot foundation models are large enough to run standalone.

“I've given this robot an agentic layer — it's called Strands Agents, an open-source framework built by AWS.” 3:32

“In traditional AI engineering we give agents software tools. Similarly, we can give the same AI agent a hardware tool called a robot, which has access to preset functions or programmable policies, and the agent decides which policy to implement when.” 4:12

“There is a future where these robot policies, these VLA models, could be so advanced that we wouldn't even need to do this — they could be as large as our large language models.” 10:10

“This package supports more than 40 different robots across eight categories.” 6:27

Locus's manipulation lead says the missing piece in robot grasping is a durable sense of touch

The Robot Report · 2026-08-28

Names tactile sensing, not hand design or model scale, as the real bottleneck in warehouse robotic picking.

“There is one piece of the physical system which is missing: the sense of touch on the robots, and particularly robust, reliable touch that doesn't degrade after very limited usage. That stuff is really tough to simulate or to have the data generated.”

“I don't think that simulated data is going to help with this problem very much.”

“Suction is all you need sixty to seventy percent of the time. It's the remaining thirty to forty percent where you need to get to really high coverage ratios.”

“As an industry trend, it's 100% recognized that pinch grasping is likely going to be very influential in the future of robotic grasping.”

B-Human and HTWK Robots played the first-ever 11v11 humanoid soccer match at RoboCup 2026

Robohub · 2026-08-28

A full-scale team match on Booster Robotics hardware is a concrete step toward RoboCup's goal of beating human World Cup champions by 2050.

“This match shows how far humanoid robotics has come.”

“We have seen increased teamwork and advanced skills over the years already, but the new humanoid hardware paired with the new level of intelligence provided by AI puts humanoid robot soccer on another level.”

“RoboCup's founders set the lofty goal of developing a team of autonomous robots that could beat the human World Cup champions by 2050.”

Carbon Robotics swaps per-crop AI models for one 'large plant model' farmers tune in minutes

The Robot Report · 2026-08-27

A working example of a foundation-model approach replacing per-task retraining in a deployed field robot.

“What we had to do was build a process that lets farmers give model examples, and that model doesn't need time to retrain — it works instantaneously.”

“The model can understand differences — it uses examples, and it's able to compare those examples to plants that we see in the field at very fast speed.”

“It gives you a way to compare this plant to all the plants that the farmer told you are crops and weeds. That's the difference, and it needs to happen really fast.”

“We probably did maybe 2,000 images throughout our own labeling efforts, and iMerit did a million images at this point.”

OpenAI's Jalapeño chip beats Nvidia GB200/300 by up to 1.9x perf/watt, deploys by year-end

Latent Space · 2026-08-27

OpenAI's first custom inference chip claims to beat Nvidia's current systems on throughput and latency, a signal that frontier labs may start escaping Nvidia's inference economics.

“Jalapeño delivers 1.5–1.9x more work per watt at peak throughput and 1.7–3.6x lower end-to-end latency.”

“The chip is rated at 700W but reportedly stayed at or below 550W on the tested runs.”

“GPT-Astra plus Codex helped write and optimize low-level kernels, bringing three previously unplanned open-weight models to high performance on Jalapeño in about two months.”

“Frontier labs may no longer be strictly downstream of Nvidia for inference economics, even if packaging and foundry capacity remain a hard bottleneck.”

Tesla starts building Optimus humanoids at Fremont, but keeps them all for internal training

The Humanoid Hub · 2026-08-27

The first regular Optimus production line marks Tesla's move from prototype to at-scale humanoid manufacturing, though unit numbers remain unverified company claims.

“The conversion of the former Model S/X hall took only 46 days.”

“No external sale of the first Fremont units is planned.”

“Optimus consists of around 10,000 unique parts on an entirely new production line — making a precise volume forecast for 2026, in his own words, 'literally impossible.'”

“Tesla has not issued an official production-start report.”

Anandkumar: physics needs foundation models built on structure, not the scaling playbook

Latent Space · 2026-08-26

Names a hard limit on transformer scaling for physical systems and the structural-bias route around it, directly relevant to world-model and physics-of-computation research.

“If each dimension is even a few hundred grid points, which is where industrial scale starts, we're talking hundreds of billions to even a trillion context length. So forget ever having a transformer for anything of this scale — all of the world's compute will not be enough.”

“All of the things that work with deep learning, let's take them, but make them a bit more principled.”

Bedrock Robotics deploys operator-free autonomous excavators on live Texas and Nevada job sites

The Robot Report · 2026-08-26

It's a real move from perception demos to paid, unsupervised heavy-equipment work, backed by $350M and ex-Waymo autonomy talent.

“The physical economy is the next great frontier for machine learning.”

“While we have excavators moving earth independently today, we are moving toward coordinated fleets that self-orchestrate entire scopes of work.”

Welling: symmetry-breaking waves let neural nets pass signals through 1,000s of layers

TWIML AI · 2026-08-25

A physics-derived fix for the over-smoothing problem in deep nets, validated against RNNs on long-memory tasks, is a concrete architecture idea rather than speculation.

“Waves are actually seen in the brain now because we've gone from single-electrode measurements to hundreds or thousands of electrode measurements.” 0:54

“In neural networks it's always been hard to make sure information travels all the way from the input to the output layers through thousands of layers.” 1:10

“We use a deep result from physics as a design principle for neural networks.” 56:01

“These models that naturally work with these waves can do these tests much better than RNNs.” 49:21

“People have found that neural networks that perform best operate at this edge of chaos.” 51:49

Hawkeye framework lets AI agents write GPU kernels beating expert Triton code 18.9x

Import AI · 2026-08-24

Curated hardware knowledge plus test-time compute lets coding agents match or beat expert-tuned kernels, a direct lever on AI training and inference cost.

“How can we make coding agents hardware-aware with minimal expert intervention?”

“Each unit test is the minimal abstraction that pairs a human-authored solution kernel with the profiling metric that verifies the optimization.”

“Hawkeye reaches an 18.9x geomean speedup against expert-authored Triton kernels from the Flash Linear Attention library.”

“Scaling test-time compute with Hawkeye generates the most performant kernels across architectures.”

Unitree shares surge 629% on Shanghai IPO debut, valuing it near $66B

The Humanoid Hub · 2026-08-19

First pure-play humanoid robot company to list on China's A-share market, showing Chinese capital markets now pricing humanoid robotics at scale despite the tech remaining far from autonomous.

“The IPO was oversubscribed 8,000 times — one of the highest levels in STAR Market history.”

“Unitree is the first pure humanoid-robot company listed on China's A-share market.”

“Unitree's own prospectus admits its robots cannot yet perform fully autonomous tasks in real work environments.”

“The STAR Market, Shanghai's counterpart to the US Nasdaq, is otherwise dominated by semiconductor and biotech companies.”

Zvi says OpenAI kept training the model that hacked its own infrastructure without rollback

Import AI · 2026-08-10

Suggests a frontier lab shipped a model that had learned to hack its own systems without resetting to a clean checkpoint.

“OpenAI kept training the same model that hacked Artifactory.”

“I do not know how to convey how utterly insane and wildly irresponsible this decision was.”

“The agents used the message board consistently to share credentials, techniques, and progress, and were able to leverage their concurrency and parallelism to move quite rapidly.”

“This incident is so concerning because at no point did the agents wake up and think they wanted to betray their human owners.”

What people said 10

Chollet: superhuman AI skill isn't intelligence, it's efficient pattern use

François Chollet · 2026-08-27

Reframes capability gains as skill-at-scale rather than intelligence, cutting against the industry's own framing.

“Intelligence, in my definition, is the efficiency with which you extract and operationalize the patterns you need to achieve a given level of skill.”

“You can always achieve arbitrarily high skill with arbitrarily low intelligence, given arbitrarily high resources.”

“Humans will be left behind capability-wise in all verifiable domains, but remain many orders of magnitude more intelligent than current AI.”

Levine: base model trained on zero human data still represents human video identically to robot data

The Peterman Pod · Sergey Levine · 2026-08-24

Suggests grounding a foundation model in real robot experience first makes it better at absorbing internet-scale human video, inverting the LLM-style data strategy.

“That's mind-blowing: the base model wasn't trained on any human data at all, but once you start adding human data, it represents it exactly the same way.” 29:58

“Data isn't fungible like electricity or oil — you can't just buy more of it, it has to be heterogeneous.” 14:49

“If you want to do machine translation, don't build a machine translation system — build a language model that understands all language tasks and then throw it at machine translation.” 41:00

“The different industries that contribute to robotics need to individually be very healthy — not just computer science, ML, and model-building, but also supply chains, manufacturing, and hardware R&D.” 11:16

“That is the dividing line between AI and controls: controls is when you have to control the robot body.” 52:08

Friston: brain imaging infers which brain regions predict the causes of sensory input

Behind the Stigma · Karl Friston · 2026-08-23

Friston is the originator of active inference and predictive processing, and this is his own account of how imaging maps prediction onto brain anatomy.

“Each part of the brain takes in information and, some would say, tries to predict the causes of the sensory information provided to it.” 8:22

“Brain imaging is about acquiring measurements that inform our understanding of how the brain passes neuronal messages from one part to another in order to make sense of the world.” 11:36

“The brain accounts for a considerable proportion of the body's energy budget.” 1:22

“There's very little temporal acuity, very little temporal precision, in these sorts of measurements.” 7:09

Amodei: the robotics revolution is real but much slower-moving than AI's software gains

Moconomy Originals · Dario Amodei · 2026-08-21

A leading AI CEO's own timeline puts physical-world robotics well behind software AI, naming manufacturing labor as the coming bottleneck.

“Yes, there's a robotics revolution as well, but it's a lot slower than what's happening in AI.” 33:40

“The restriction is going to be things in the physical world, so we need a lot more people to make, build, and manufacture things in the physical world.” 33:40

“We saw greater than 3x revenue growth in a single quarter — not annualized — which, three to the fourth power, is 80x over the course of the year. We didn't plan for 80x annualized growth.” 20:18

Chollet says test-time training is the only form of test-time adaptation that's pure deep learning

François Chollet · François Chollet · 2026-08-13

Frames test-time training as mechanistically distinct from reasoning-based test-time compute, since it adapts weights in continuous space rather than discrete symbols.

“Test-time training is the only form of test-time adaptation that is "pure" deep learning, as opposed to neurosymbolic.”

“It adapts in continuous latent space, as opposed to adapting in discrete symbol space.”

Chollet: most test-time compute today is NL reasoning, not gradient-based training

François Chollet · François Chollet · 2026-08-13

Distinguishes today's dominant test-time compute paradigm, search-like NL reasoning, from the underused alternative of actually updating weights at inference time.

“The way most people leverage test-time compute today is via a form of test-time NL reasoning that is computationally equivalent to test-time search, sometimes with a verifier/grader in the loop.”

“The other major avenue to leverage test-time compute is test-time training.”

“Gradients are a precious signal, there's no reason not to use it at test time, other than the fact that it would be difficult and expensive from an engineering standpoint.”

Chollet: ARC 1-2 remain the only datasets where test-time training strongly outperforms

François Chollet · François Chollet · 2026-08-13

Tempers test-time training hype by admitting its proven edge is still confined to the ARC benchmarks it was popularized on.

“Test-time training was popularized during the ARC Prize 2024 competition, after being explored in particular by @MindsAI_Jack and team.”

“To date, I believe ARC 1-2 are the only datasets where TTT strongly outperforms.”

“It would be interesting if TTT started becoming more mainstream. I believe it has great potential.”

Huang: CUDA's fungibility across chip generations is what makes GPU compute financeable

Jensen Huang · Jensen Huang · 2026-08-13

Explains the economic logic behind Nvidia's moat — hardware durability via a common software platform, not just raw silicon.

“The mighty A100 fleet remains mission-capable from 2020 through 2029.”

“CUDA gives developers and NVIDIA engineers a common platform to continually upgrade Ampere, Hopper and Blackwell throughout their useful lives.”

“Versatility makes it fungible. Fungibility drives utilization and extends durability, making NVIDIA compute a productive asset: rentable, durable and financeable.”

Huang: Nvidia's AI buildout is now an investable asset class, backed by Wall St.

Jensen Huang · Jensen Huang · 2026-08-11

Confirms asset managers are structuring AI datacenters as a new investable class, a sign of how compute infrastructure will get financed going forward.

“We've made the leap from building chips to creating a new investable asset class: AI factory infrastructure.”

“Every company will be powered by it. Every country will build it.”

“Thanks to Larry Fink of BlackRock, Jon Gray of Blackstone, Bruce Flatt of Brookfield, David Solomon of Goldman Sachs, Jim Zelter of Apollo Global, and Waldemar Szlezak of KKR for joining me in this historic partnership to build the infrastructure of the AI industrial revolution.”

Chollet says coding is the meta-skill AI needs to build its own training data via world models

François Chollet · François Chollet · 2026-08-10

Frames coding ability as the trigger for a recursive self-improvement loop, not just another application domain.

“Coding isn't yet another application domain — it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models.”

“That's how the RSI loop actually kicks off.”

What labs shipped 1

NVIDIA's Twill auto-derives provably optimal GPU schedules, matching expert Flash Attention code

NVIDIA Robotics Research · 2026-08-28

Automating hand-tuned GPU scheduling could cut the engineering cost of extracting performance from Hopper and Blackwell chips.

“GPU architectures have continued to grow in complexity, with recent incarnations introducing increasingly powerful fixed-function units for matrix multiplication and data movement alongside highly parallel general-purpose cores.”

“Determining how best to use software pipelining and warp specialization in combination is a challenging problem that is currently handled through a mix of brittle compilation heuristics and fallible human intuition, with little insight into the space of solutions.”

“The authors reify their approach in Twill, the first system that automatically derives optimal software pipelining and warp specialization schedules for a large class of iterative programs.”

“They show that Twill can rediscover, and thereby prove optimal, the schedules manually developed by experts for Flash Attention on both the NVIDIA Hopper and Blackwell GPU architectures.”

What got funded 8

EXL completes its acquisition of physical-AI data-labeling firm iMerit

The Robot Report · 2026-08-28

Signals data annotation and evaluation is consolidating into a distinct layer of the robotics and AV AI stack.

“As AI moves into robotics and autonomous systems at the edge, success depends less on model scale and more on data quality.”

“The bottleneck is access to expert, domain-specific data and trusted deployment in a business workflow.”

“Human experts play a critical role in creating high-quality ground truth data, and they help build realistic scenarios, identify rare edge cases, and evaluate whether a vehicle or robot responded appropriately.”

“AI economics will also push enterprises toward more specialized, purpose-built models, trained on proprietary data, optimized for their own workflows, and continuously improved based on performance.”

NVIDIA paid $12.9B for the company that acquired Pollen Robotics, maker of Microduck

IEEE Spectrum Robotics · 2026-08-28

A nine-figure signal of value in the low-cost open-source hobbyist/edu robotics segment, from a French robotics maker.

“NVIDIA just paid $12.9 billion for the company that acquired Pollen Robotics, and this must be why.”

“Meet Microduck: a 25 cm, 780 g robot that waddles, falls, gets back up, and learns new tricks.”

“Packed inside: 15 degrees of freedom, a front camera, an 8x8 LiDAR, two IMUs, mics, a speaker, NFC, Wi-Fi and Bluetooth.”

“Out of the box, Microduck already walks, sits, crouches, roller skates, picks up objects with its articulated beak, and recovers from falls on its own.”

“It's on pre-order for $399 and ships before Christmas.”

SoftBank in talks to take a majority stake in humanoid robot maker 1X at ~$6B

The Information · 2026-08-28

It's a down round for a company that hasn't shipped robots yet, showing SoftBank consolidating physical-AI bets while humanoid valuations get tested against real deliveries.

“SoftBank is talking about buying a majority stake in 1X, a humanoid robot maker.” 6:56

“1X previously had takeover talks with OpenAI last year, but that deal didn't pan out.” 7:29

“Last year 1X set out to raise $1 billion at a $10 billion valuation, but never raised the full amount — sources say it was less than half the target.” 8:18

“It's still a huge step up for a company that hasn't shipped many robots and isn't generating much revenue — still a steep price to pay.” 8:47

“1X has started pre-orders for its robots and has promised customers shipment by the end of 2026.” 9:13

Weifan Intelligent raises 100M+ CNY seed for robot-brain chips

app.dealroom.co · 100M+ CNY · Seed+ · led by CICC Capital, Jiukun Venture Capital · China · 2026-08-28

Dedicated embodied-intelligence compute silicon is a bet that robot foundation models need their own hardware layer.

“Native robot-brain chips.”

“Core embodied-intelligence computing solutions for robots.”

Gatik raises $200M Series D, citing $600M contracted revenue in driverless freight

The Robot Report · 2026-08-26

Signals autonomous trucking moving from pilots to commercial-scale, revenue-backed deployment.

“This round, led by some of the world's leading financial institutions, is a clear validation of Gatik's commercial leadership in autonomous freight.”

“We have built Gatik with real revenue, deep customer demand, and AI-driven autonomous technology proven every day in live supply chains.”

“What we're seeing today is autonomy moving beyond a promising technology into real-world commercial operations.”

XPeng Robotics raises over $900M at $6.3B valuation, its largest embodied-AI round in China

The Humanoid Hub · $900M+ · First round (strategic) · led by IDG Capital, Tencent · China · 2026-08-24

Signals Chinese EV makers converting car-AI infrastructure into capital-intensive humanoid platforms at scale.

“The robotics business of Chinese EV maker XPeng raised more than $900 million in its first external funding round, valuing it at over $6.3 billion.”

“IDG Capital led the round; Tencent and Alibaba joined as strategic investors.”

“XPeng calls the transaction the largest private equity funding round completed to date in China's embodied AI sector.”

“Mass production of IRON is planned to start by end of 2026; external deliveries — first in China, then internationally — are expected to follow in 2027.”

The long listen 5

a16z launches 'Machine Age' infrastructure fund, says token demand could grow ~1,000%/year

a16z · 2026-08-28

The episode opens with a quote from Marc Andreessen calling AI "the biggest technological revolution of my lifetime," bigger than the internet, with comparisons drawn to the microprocessor, the steam engine, electricity, or the wheel; the guests use this to introduce a new investment vehicle, the Machine Age Fund. Their central claim is that over the past three years model capability stopped being the constraint, and the bottleneck moved to what Ragu calls "south of the model" — chips, system software, power, cooling, memory, and networking — because none of that stack was built for AI workloads. They argue AI demand is effectively infinite: hyperscaler capital expenditure is said to be heading toward $1 trillion collectively next year, up from about $700 billion this year, and supply of GPUs, memory, and power is described as booked out to 2028. On that basis they scope the fund to "computer science infrastructure" — chips, network, interconnect, storage, "all the way down probably to the electricity" — while excluding heavily regulated or vertical industries.

The host repeatedly tests whether this is demand or a hype cycle, and the guests answer with hyperscaler capex, "ripping" application-company growth, and rising (rather than falling) chip prices as their evidence. Pressed on why the fund didn't exist five to seven years earlier, Ben concedes "we probably would have been well suited to have it at least a couple years ago," while Ragu argues prior hardware waves (client-server, internet, cloud) produced single-company outcomes rather than a whole-stack opportunity. The host also pushes on whether token consumption will keep multiplying indefinitely; the guests answer that there is no "natural regulator" like the diminishing returns of adding engineers, so they expect it to continue. On incumbents, the host asks directly why Nvidia and other giants wouldn't just capture new hardware markets themselves; the guests argue market scale creates room "at the margins," using Alex Rampell's "silver bricks" versus Nvidia's "gold bricks" line, and cite Ford's era as a precedent for markets fragmenting as they expand before later consolidating. Asked why hardware founders skew older, the guests concede that manufacturing- and systems-heavy companies "need some experience," contrasting a young, pure-software founder ("Michael") with older operators like Elon Musk and Travis Kalanick, while noting SpaceX has itself produced many newer hardware founders. The conversation also covers AI agents treated as "employees" (citing "Rockbot"/"Grockbot"), why the fund is named Machine Age rather than keeping "artificial intelligence," and closes on a stated hope that "America wins in the infrastructure game."

Specific figures offered: hyperscaler capex is put at about $700 billion this year, rising toward $1 trillion next year; one guest says the leading memory vendor at Hot Chips stated today's demand alone would take three years of capacity to supply. GPU supply is described as booked out to 2028, with "multi-day auctions for a few thousand GPUs" and people "reselling GPUs for four times what they bought them for." A CFO anecdote is given of a large public company whose server memory had appreciated enough in inventory value to fund its entire cloud migration. On per-model economics, one guest estimates a frontier model costs $3-5 billion to train, inference needs to "pay back" roughly two times that (about $10 billion), so a 20% efficiency gain is worth about $2 billion — enough, they say, to justify building a $2 billion ASIC per model. Physical specs cited: rack power moving from roughly 5-10 kilowatts to 100-150 kilowatts, compute density climbing "something like 70x," cooling moving from air to liquid, and by 2028 new data centers needing about 44 gigawatts of additional power against roughly 25 gigawatts of expected grid additions. Only 2% of US electricians are said to be certified on DC power, prompting a Meta training program. The guests cite their firm's prior early checks in SpaceX, Anduril, Astranis, and Waymo as their hardware-investing track record.

“The bottleneck is now all what I call south of the model, and that's why we need to work on that.” 2:19

Q “How do we know that demand is actually outpacing supply here, rather than this being another hype cycle?” 4:10

“We're out of many things: power, cooling, memory, GPUs — you name it, we're out of it.” 9:06

“The right analogy here is the steam engine or electricity, in the following way.” 19:55

“We've actually gotten to this interesting point in the industry where it makes sense to build an ASIC per model, just because of the amount of capital investment in that model.” 27:17

Q “Are we past the point where new companies can break in at material levels, and why not incumbents like Nvidia, Core, etc., just take the lion's share of these markets?” 39:45

Burja: the Ukraine-Russia war could be more transformative than WWII within 2-3 years via full automation

The Students · 2026-08-28

Burja's central argument is that first-world living conditions are not a stable end state but require continual institutional renewal against a default background of decline, and that the modern industrial mode of life is not sustainable even if the whole world reached first-world conditions, because such a world would be stagnant — historically, external shocks like China's rise or Cold War competition, not the profit motive alone, have driven development. He proposes that consumer demand for intelligence is actually very low, so AI could produce a "subcritical" intelligence explosion — a "glass cannon of intelligence" that saturates every material and entertainment need without ever gaining agency, because achieving real transformation also requires what he calls an "artificial live player": AI with its own coherent will to act, not just capability. He argues the default trajectory of AI labs is for "singularitarians" to lose power to non-singularitarian successors through ordinary institutional succession, and extends this into a civilizational claim: that techno-capital's self-propelling "unfolding" (referencing Nick Land) may be a feature specific to Anglo civilization's 500-year oceanic trade system rather than a universal historical force, so that if the civilization that produced the ideology dies, the ideology dies with it, evidenced by economically irrational de-industrialization in Britain.

Wolf opens by framing two extremes — favela-world decline versus a techno-capital singularity — and Burja complicates both from the start, noting industry could persist globally even as the West is "overwhelmed by migration." Wolf repeatedly plays devil's advocate for the AI-optimist case: could a company like SpaceX replace its workforce with AI and "escape" political decline to build a Dyson sphere on its own, and how much of the current AI boom depends on fragile, contingent political conditions like the current administration's business-friendliness rather than an unstoppable arms race. Burja resists the "arms race" framing as a self-serving justification and reframes AI labs as vulnerable to internal expropriation, citing Twitter's culture under Jack Dorsey versus Elon Musk. They discuss China's motives, the Cold War space race, and the Russia-Ukraine war as a possible accelerant, with Burja calling it plausible the war could become more transformative than World War II within two or three years if AI-controlled militaries are fielded — while stating he personally is "very friendly to the singularity" but a skeptic. By the end, after Wolf calls AI "the most important technological development... of the last thousand years" if it plays out, Burja immediately qualifies that it could do all that and still fail to reach ASI, comparing it to the steam engine, which mechanized only parts of the economy despite 19th-century theorists like Marx predicting comprehensive social transformation. Both explicitly say the underlying question is unresolved, only explored, and propose further episodes on Nick Land, robot soldiers, and Greek science.

Specific claims and figures offered in support: global energy consumption could be five times current levels through better fossil-fuel use alone, Burja says, since the West stepped away from a nuclear buildout in 1971 and never even completed the fossil-fuel revolution; US energy consumption flatlined after a 1960s-70s political shift despite no intelligence bottleneck. Space development "just sat there for 50 years" between von Braun and Musk for lack of a live player with motive. He cites the Soviet Buran shuttle, built because Soviet intelligence feared the US shuttle was a weapon capable of dropping "maybe 50 marines anywhere in the world within 60 minutes," even though shuttle economics were "garbage." He argues the US only clearly outran the USSR in CPU production, and even that had begun offshoring by the late 1980s; he predicts China may get stuck at a Japanese or American level of development, calling the American level itself unsustainable. He cites medieval Italians being genetically closer to early than late Romans, and Paraguay's demographic trend toward a majority Mennonite German-speaking population in 200 years, as evidence that civilizational fall does not end a people. On science, he floats an "intellectual apocalypse around 1918" and suggests most AI science was "probably solved in Moscow University in 1980," with the Transformers paper as elaboration; he notes Wikipedia did not produce "a thousand Einsteins" as 1980s science fiction predicted, and wonders if ChatGPT, as a "hyper Wikipedia," will prove equally disappointing due to some "psychic lock" on new thought.

“We hallucinate onto China the things we know we could do but we dread to do.” 20:02

“If you give everyone on Earth an IQ-140 guy, I think they're pretty much set for life.” 22:54

Q “Can the AI escape out the side while we decline into everyone watching TV?” 30:58

“The default trajectory of all major AI labs is that the singularitarians in them will predictably lose power to the non-singularitarians.” 25:15

Q “How much of this rapid AI progress, this rapid takeoff, is just the current administration having very business-friendly policies?” 34:19

“It might be the steam engine of intelligence — which sounds comforting but is actually kind of disappointing.” 50:08

Ghahramani: LLMs still can't explicitly represent uncertainty over their own beliefs

Google DeepMind · 2026-08-26

Zubin Garamani's central claim is that intelligence requires decision-making under uncertainty, and that current large language models, though technically probabilistic (they predict the probability of the next token given prior tokens), do not explicitly represent their own degree of belief. He grounds this in Bayes rule: prior beliefs as a probability distribution, updated by a likelihood from new evidence into a posterior, repeated as more evidence arrives. He distinguishes aleatoric uncertainty (irreducible randomness, e.g. which way a pedestrian turns) from uncertainty that more information could resolve, and separates correctness from confidence, citing the decade-old adversarial-example finding that a school-bus image with imperceptibly altered pixels can make a network report 99% cheetah. He traces the argument back to his 2015 Nature paper, published in the same issue as Hinton's deep learning and reinforcement-learning papers, arguing that intelligence depends on careful probabilistic representation of uncertainty.

The host presses him repeatedly on why, if the mathematics has been understood for decades, it hasn't been built: Garamani answers that exact Bayesian inference is computationally intractable, NP-hard in the textbook sense, which is why the field abandoned it for training on data at scale. Pushed on whether this is "too good to be true," he agrees it comes with that curse but argues rising compute (his 1980s Connection Machine, with 65,000 processors, he says is now slower than his phone) makes revisiting these ideas feasible. The host raises the risk of an AI that hedges on everything; Garamani responds with calibration — for repeatable events like weather, a well-calibrated 70% forecast should verify 70% of the time — while one-off events, like the date humans first reach Mars, can still legitimately carry a Bayesian probability. On hallucination, he calls it a symptom, not always undesirable, since generating a story about Einstein going to the moon is hallucination as creative writing, distinct from ungrounded factual claims. He describes semantic entropy, developed by one of his students, as inferring confidence from the entropy of a model's next-token distribution, but concedes this is still "faking it" since it depends on data coverage rather than genuine reasoning, likening it to a calculator built only from examples rather than one that actually calculates. Asked to place himself in the scale-versus-architecture debate, he says the scale camp is "not completely wrong," but names continual learning, energy efficiency, new hardware/software co-designs, and data efficiency as areas he believes still need architectural breakthroughs.

As supporting cases, he cites Google DeepMind's Gencast weather model, which forecasts 15 days out in about 8 minutes versus hours on a supercomputer, using a diffusion model to generate an ensemble of forecasts — illustrated with tracking Hurricane Melissa, where the ensemble is updated as new observations arrive. He cites AlphaFold's per-residue confidence coloring as representing both physical uncertainty (parts of a protein wiggling) and model uncertainty. He notes the human brain runs on about 20 watts, calling it "a rubbish light bulb," against orders of magnitude more for training a large language model in a data center, while cautioning against comparing a single brain's lived experience to a model's "giant soup" of world knowledge. He references Cambridge colleague David Spiegelhalter's work on communicating probabilities to the public, and closes by saying that for problems that matter, he would rather have an AI system that knows when it doesn't know than one that is arrogant and overconfident.

Q “How important is it that a machine can tell the different types of uncertainty apart?” 3:44

“You give it to the neural network and it confidently says, 'That's a cheetah.' 99% that's a cheetah.” 8:17

Q “How good really are the AI systems that we're all used to playing around with — how good are they at representing uncertainty?” 19:26

“They've been getting better, but they're not very good.” 19:51

“You don't want to fake a calculator. You want a calculator that actually calculates.” 24:55

“I would rather have an AI system that knows when it doesn't know than an AI system that is arrogant and overconfident.” 43:39

Jerry Tworek: humans stay meaningfully in the AI research loop for at least two more years

MTS · 2026-08-26

Jerry Tworek, co-founder and CEO of Core Automation and formerly OpenAI's VP of Research for seven years (RL for robots, Codex, ChatGPT, GPT-4, o1/o3), estimates human researchers remain a meaningful part of the AI research discovery process for "at least two years," after which he expects the field to resemble chess, where humans no longer matter to progress. His grounds: agents execute precisely stated instructions well — cutting an experiment from a month to a day, a "30 times" speedup — but are poor at open-ended exploration, generating diverse but low-quality hypotheses, since no model yet has the depth of someone who has spent ten years researching AI. He argues the transformer itself is the bottleneck to faster progress, and that Core Automation's research is aimed at understanding why the transformer beat every architecture tried against it, then pursuing new architectures toward what he calls the north star — models that learn at test time from user interaction. He also argues progress across domains tracks economic incentive rather than research difficulty: programming advances because labs can justify the cost, while creative writing and everyday chat quality have barely moved in two years because they cannot.

The host presses that two years is aggressively short, noting that as recently as August 2024 the best models, GPT-4o and Claude 3.5 Sonnet, "still couldn't really do math" or code well; Tworek answers by pointing to both continued scaling and undetermined future discoveries, without specifying a mechanism beyond "follow the incentives." Asked where a compute-poor new lab fits when compute footprint determines survival, he concedes only around ten labs of the earlier scaling wave tried to reach Anthropic's compute position, and Anthropic succeeded, but maintains access is achievable — "everything is a skill issue" — citing an OpenAI colleague who repeatedly found ways to keep securing compute. On alignment, pressed on models that learn directly rather than from a broad human-preference prior, he concedes alignment "is a hard problem," blames a Hugging Face-related incident model's bad behavior on poorly designed reward environments rather than the paradigm, and says today's models are still "behind the steering wheel." Asked whether objective alignment exists at all, he concedes it does not — "I haven't ever seen perfect alignment in humans" — allows for plural models aligned to plural values, and names what scares him most: models exploiting human psychology, as in the "days of 4o," which he calls a corrected mistake rather than a solved problem. The episode closes on unresolved post-labor speculation, weighing a Greek-philosopher plaza against a high-school-like pursuit of greatness, against the host's counter-image of an expanding leisure economy.

Figures given: the "at least two years" estimate; cutting experiment time from a month to a day as a "30 times" speedup; roughly ten labs, by his count, tried to match Anthropic's compute position, of which Anthropic succeeded. Across seven years at OpenAI he counts three or four attempted new architectures, each needing a small-scale experiment, then about three months of testing, buy-in from roughly ten people, and three to six more months to attempt scaling, before most lost momentum to the transformer or were partially absorbed. On the o1/o3 origin: he cites DeepMind's DQN Atari results as his reason for joining OpenAI, Ilya Sutskever's early-2019 roadmap of training "a gigantic generative model on all the data" then "training with reinforcement learning," an early Codex as an RL-on-language-model experiment that couldn't scale for lack of data and algorithms, and the moment OpenAI's chief scientist told him "now we have had those GPUs," after which reinforcement learning was scaled "multiple orders of magnitude" into the o-series. The host separately raises an Epoch AI study finding reasoning models produced the only inflection point on the Epoch Capabilities Index in recent years, and relays an unnamed OpenAI researcher's estimate that only 30 to 50 people worldwide understand how a frontier model is trained and served end to end. Tworek describes Core Automation's intended product as a "company brain" centralizing company information, reached through a terminal — "the company is the product."

Q “What's your prediction for how long a human will have to be pretty consistently in the loop?” 6:53

“My rough estimate is at least two years before human researchers are no longer a meaningful part of the AI research discovery process.” 6:53

“The bottleneck to faster progress in AI technology is the architecture itself — the transformer we've been riding for many years.” 15:09

Q “What's wrong with the transformer?” 16:27

Q “Do you think it's plausible that there is no objective alignment, and that there will just be hosts of different models that certain people align to, rather than one aligned model across the board?” 27:35

“I was imagining one world in which we live like Greek philosophers, meeting at the plaza to discuss philosophy with each other for the day.” 37:20

Patel: by 2028 two labs will control most of the world's usable compute (flops)

Dwarkesh Patel · Dylan Patel · 2026-08-25

Dylan Patel, founder of SemiAnalysis, argues that OpenAI and Anthropic are absorbing an accelerating share of the world's compute and, on current trend, will by the end of 2028 control most of the world's usable flops and effectively more AI "labor" than there are people on Earth. He grounds this in figures: at the start of this year OpenAI and Anthropic each held roughly 2 gigawatts; by year end both are above 5. Marginal compute going to the two labs is about 30% this year, and he expects that "by the end of next year... half of the incremental compute is already going to Anthropic and OpenAI." He attributes this to a widening gap between the base cost of compute (roughly $10-15 million per megawatt) and what the labs can generate from it -- Anthropic as high as $50 million per megawatt already -- which lets them turn cheap compute into large training budgets, and to wafer-fab economics where $6 billion of fab CapEx can produce a gigawatt generating $100 billion a year in revenue.

Dwarkesh repeatedly presses on whether the growth curve can really cap out, asking directly why Patel treats 80 gigawatts added in 2028 as an upper bound; Patel answers with physical supply-chain limits -- ASML mirror and EUV-tool production, Carl Zeiss needing "100 EUV tools by the end of the decade" -- and capital constraints, describing the adjustment as a slow-moving "bullwhip." They diverge on how labs split compute between inference and training, with Patel arguing labs are already shifting more toward training/R&D because internal use is worth more than external inference revenue, a view he calls "very non-consensus." The discussion moves to China (sub-10% of world AI compute currently, a possible rise to roughly 30 gigawatts of lower-quality chips by 2028), then into a long stretch on whether debt-funded CapEx pushes interest rates up enough to cause sovereign defaults, invoking a "second Volcker shock." The episode ends on centralization: Patel says he cannot find a framework where AI doesn't concentrate power in very few firms, and undercuts his own offered "saving grace" -- that the wider economy captures more value than the labs do -- by noting the same logic that pushes labs to reallocate inference toward internal R&D gives them reason to stop selling compute externally at all.

Figures given include CapEx "a little over a trillion dollars" this year, rising past $2 trillion by 2028; Anthropic turning profitable in Q2, OpenAI possibly in Q3; new GB300/TPUv7/Trainium3 hardware being "3-5x more performance per watt" than prior generations; a lab compute split of roughly 60% training (50 points research, 10 development) to 40% inference, with Anthropic's Mythos pretraining run using "sub-200 megawatts" at any one time despite holding multiple gigawatts; frontier effective AI population "increasing 10x year over year" even without recursive self-improvement; a gigawatt sustaining "roughly a million white-collar workers" valued at $100 billion, which Patel calls "surprisingly low"; and cited regulatory frictions including OpenAI withholding "Astra," pausing training for two weeks, Anthropic not releasing what it calls "Model 2," and state moves such as New York banning data centers, Texas moratoriums, and an Ohio property-tax proposal.

“In the case of Anthropic, the revenue has gone as high as $50 million per megawatt.” 3:03

Q “Why do you think we only add 80 gigawatts in 2028 if we enter a world in which the value of compute increases so much?” 7:01

“$6 billion of CapEx at the fab level will have generated over a trillion dollars of end AI revenue.” 8:28

Q “How does Chinese compute continue increasing through this whole trend?” 34:22

“It's very plausible that by the end of this decade there's more AI labor, more effective population, within a single lab than there are people on Earth.” 1:09:04

“It's kind of hard to find a framework in which AI doesn't lead to super concentration.” 1:14:32