CLANKERS WEEKLY

Week of 2026-09-21 · 40 items

Robotics firms raised $4.87 billion in August, about half in China, and AMI Labs closed a $1.03B seed while UBTech opened a Liuzhou plant sized for 10,000 humanoids a year, against about 7,000 humanoids sold worldwide in 2025. European hardware founders now compete on manufacturing scale: 1X targets 50,000 humanoids next year, and SoftBank agrees to acquire Hyundai's Robotics and AI Institute pending CFIUS review.

What happened 20

IFR counts about 7,000 humanoid robots sold in 2025, many bought to generate AI training data

The Humanoid Hub · 2026-09-21

It is the first industry-wide count of humanoid sales, and it shows a small base in which a considerable share went to research and training-data collection.

“This is the first systematic, cross-industry survey of its kind and for the first time provides a reliable baseline for the still young product class.”

“Legs are explicitly not required: wheeled chassis or traction units are permitted.”

“The IFR stresses that a considerable share of the 7,000 units was not used for productive work. Many units were acquired by research institutions and companies to generate training data for AI models.”

“The pilot projects typically run with single-digit to low double-digit numbers of robots per plant, a sign that the transition from proof of concept to scalable deployment has not yet been made.”

“Chinese vendors (AgiBot, Unitree, UBTech) dominate with a global market share of over 90%.”

Rui Ma says Huawei is offsetting weaker chips with larger clusters, aiming for a million-chip system

The Information · 2026-09-20

Scaling cluster size is the route Chinese labs are taking around export controls on chips, and it bears on whether they reach the frontier.

“The chips are absolutely the limitation. I think it is, globally really, but for China specifically because of these very stringent export controls.” 7:56

“What they have announced for next year is a 4,000 chip cluster as a pod, and their goal is to connect that into a million chip cluster. And in fact, they said that they have a 256,000 chip cluster in deployment right now.” 8:46

“Chinese model companies are already training 1 trillion parameter models. They are not the frontier models just yet, but that might happen with the next generation or two with Huawei chips. We do not know.” 9:39

“We do not see this existential risk of rogue AI agent swarms wiping out humanity because we do not fit in their goals as an immediate risk.” 4:38

“I do not see anyone actively investing in Chinese companies at scale, and part of it is because different capital exits. Chinese companies are increasingly exiting in Chinese markets.” 10:36

Wulff says egocentric capture with tactile gloves scales where teleoperation does not

The Robotics Stack · 2026-09-20

Robot labs ask for 100,000 to a million hours of human data a month, and Wulff argues that cheap tactile gloves worn by workers in real jobs are the only route to that volume.

“A lot of these current robots on the market are trained with video data, or they have an LLM in the background that helps them do all these actions, but a lot of them are still missing that touch layer.” 1:13

“For context, when these labs ask for data, they ask for around 100,000 to about a million hours of data every month.” 14:48

“You need 10 to 100 times that to train these humanoids, and to have people capture that much data is just impractical.” 30:35

“Going from a higher number of channels to a lower number is always easier than the other way around, because you have to make up all the missing data.” 28:44

“The labs ask for high quality data, but often it is very undefined, so they want you to say this is high quality data, and if you can convince them of that, then they get excited about that.” 23:39

Kiely says TurboQuant halves KV cache bits but cuts decode speed by over half, so Baseten skips it

AI Engineer · 2026-09-19

Baseten reports which new inference techniques work in production: DFlash speculative decoding gives over 3x, while TurboQuant suits local inference and not data centers.

“Turns out that you need to do additional computation in the forward pass to account for this during decode, and it cuts TPS by more than half, and that's just an unacceptable trade-off for a lot of the production use cases.” 7:20

“Dlash models might be two or four times slower to run, but they're going to predict eight or 16 tokens at once in that window, while Eagle is only doing one at a time.” 13:54

“If you are continuously retraining on those prompts and responses in your live system, you can see a 20% to even 2x improvement in your token acceptance rates.” 16:10

“Many optimizations for inference come from a dedicated training process. And so the lines between training and inference are getting blurrier and blurrier.” 4:04

“Still is a perceivable bottleneck that takes a fixed set of loan query vectors, crossends it against the full KV cache, and produces a set of compact keys and values in a single forward pass.” 11:41

Lambert argues scaling laws' exponential costs limit recursive self-improvement to efficiency gains

Interconnects · 2026-09-19

The essay gives a reasoned moderate position on whether frontier labs' internal automation produces an intelligence explosion, which bears on compute demand and progress timelines.

“Automatable research is too narrow to achieve a massive net acceleration in progress, in the face of scaling laws’ exponential costs.”

“Still, I cannot get past the fact that all of our scaling laws show that you need exponential compute and resources to make linear improvements in intelligence.”

“RSI is much more helpful at efficiency rather than expanding peak intelligence.”

“Our current techniques work and let us solve problems we know how to state, but they don’t result in a magical level of generalization to unknown, harder problems in most partially verifiable domains.”

“The biggest takeoff in automation within the labs is in tasks like software engineering, monitoring logs, managing planned experiments, and other fairly routine (but not always easy) tasks.”

Misili says AI compute is financed with debt, at about 6% for Microsoft use and 9% for weaker customers

The Information · 2026-09-19

Financing terms now decide who can reserve GPU capacity, so hardware and compute startups face upfront payment demands and borrowing costs tied to the credit quality of their customers.

“So much of the AI infrastructure boom is being built with borrowed money.” 48:46

“A lot of what you're seeing is cloud providers borrowing to buy the GPUs that customers then use.” 49:16

“If we're borrowing to pay for GPUs that Microsoft's using, we're paying roughly 6%. If it's a non-investment grade customer, more like 9%.” 50:39

“Getting in line right now increasingly means making a big financial commitment. So providers want to see long-term contracts.” 49:47

“A whole another bucket is the GPUs, and at the end of a contract, what are they really going to be worth? Which is a huge open question right now.” 54:31

Neptune Medical gets FDA clearance for Triton 1, a robotic system for colonoscopy

The Robot Report · 2026-09-18

It is a regulatory clearance for a robotic endoscope, backed by first-in-human data with a 54.2% adenoma detection rate against a 35% industry threshold.

“The study showed that the Triton 1 System met its primary endpoints, with no adverse events and 100% cecal intubation.”

“The robotic system also demonstrated a high adenoma detection rate (ADR) of 54.2%, compared with the industry threshold of 35%, the company reported.”

“It was also effective in creating a highly favorable ergonomic environment for endoscopists, with a 67% lower average burden compared to manual colonoscopy (based on the NASA-TLX score).”

“In 2024, Neptune Medical raised $97 million in Series D funding and launched its subsidiary Jupiter Endovascular Inc.”

Icarus Robotics tested its ISS-bound JOY robot on four parabolic flights ahead of a January NASA handover

The Robot Report · 2026-09-18

The tests validated arm control, state estimation and flight control for a manipulation-capable free flyer running on a Jetson Thor, and the suspension of the only U.S. parabolic flight operator is pushing American hardware teams to Canada.

“The arms themselves we were the most happy with. It was one of the biggest unknowns and one of the hardest things to model in the simulator.”

“So, what we were doing is taking these pre-recorded teleop trajectories and pre-recorded set points. We were able to go do these terrestrially in our office, take that data, and then run the same thing in microgravity, benchmark it, and compare those two things.”

“Icarus plans to teleoperate its robots when they arrive at the ISS and develop an autonomy system over time as it gathers data in space.”

“Where they use cell phone compute, we’re using a T5000 from NVIDIA Jetson Thor.”

“So, American companies building hardware for microgravity now have to leave the country to prove it works.”

Ermon says diffusion LLMs map onto GPUs at inference, where autoregressive decoding is memory bound

No Priors · 2026-09-18

If diffusion language models match autoregressive quality, they give custom-chip speed on ordinary Nvidia GPUs, which changes the compute needed per token served.

“You cannot generate the 10th token until you've generated everything that comes before it. That kind of workload does not map well to GPUs. That kind of workload is extremely memory bound.” 8:51

“The workload that we have at inference time in a diffusion model is basically very similar to the workload you have for training, where you're processing many tokens at the same time in parallel.” 9:41

“Even if you think about RL post-training, a lot of the bottleneck is generating rollouts.” 10:32

“They can essentially get the same speed as what you would get if you were to run an autoregressive model on custom hardware.” 18:04

“There is a decent amount of evidence in the academic literature, for example, that diffusion-based models are more data efficient compared to autoregressive models.” 28:03

“There is between 20 and 30% where latency is really important.” 29:45

Meeks says rate hikes raise borrowing costs for AI infrastructure builders and lower their valuations

The Information · 2026-09-17

Neoclouds and colocation firms finance data centre build-out with debt at subprime credit ratings, and Meeks argues that higher revenue per megawatt offsets the higher interest but that capex growth will slow.

“Most of these companies, the neoclouds and the AI colocation companies, are subprime credit rated.” 0:53

“They are getting much better pricing on their deals: two or three times what they were getting a year or two ago in revenue per megawatt.” 1:26

“Rising interest rates equal a rising discount rate for these cash flows, which really hammers the stock valuations.” 1:57

“Today 90% of the AI infrastructure building has been with four or five of these hyperscalers that have very strong balance sheets, all investment grade.” 5:09

“I expect it to go fast and furious at least through 2028, maybe through the end of the decade.” 5:42

Ousterhout says Homa cuts 99th-percentile latency for short messages about 13x against TCP

AI Engineer · 2026-09-17

Inference and agentic workloads shorten GPU compute phases to milliseconds, so the tail latency of small synchronisation messages leaves GPUs idle, and TCP and RDMA are poorly suited to that traffic.

“Whereas the workloads used to be completely dominated by large transfers, where throughput is the key metric that matters, we're seeing more and more smaller transfers where the latency is crucial.” 1:22

“But now with agentic workloads, where you're trying to pump out tokens relatively rapidly at a regular rate, the periods of computation are getting down into sort of the millisecond time scale.” 4:55

“And if it also takes milliseconds to do that synchronization, then you're wasting a significant fraction of your GPU resources waiting for the synchronization to occur.” 5:23

“These systems tend to never stabilize. They're constantly oscillating between sending too much and sending too little.” 9:12

“As soon as a receiver gets the first packet of a message, it knows exactly how much more data the sender wants to send.” 12:52

“With TCP it's more than a millisecond tail latency. Homa is less than 100 microseconds, about 13 times faster.” 16:51

Arm launches a physical AI partner program of 80+ companies and a six-level robot capability scale

The Robot Report · 2026-09-17

Arm is positioning its chip architecture as the shared basis for robot compute, drawing on its position in automotive.

“Arm Holdings PLC recently released Arm Total Design for Physical AI, a program that brings together more than 80 companies spanning the physical AI technology stack.”

“Robotics is a fragmented industry, bringing together people from a range of disciplines and industries.”

“To start, Arm has identified six levels of capability, from zero to five. On one end, you have robots that can react to stimuli and execute commands, but must adhere to fixed rules. On the other end, you have robots with evolving behavior that learns and self-optimizes.”

“From that experience, the company saw that robotics and automotive manufacturers require similar safety criticality, real-time computing, computer vision capabilities, and more.”

Tutor Intelligence CEO Greenstein says customers pay $14 to $18 an hour to use its robots

The Information · 2026-09-17

The hourly rate is a public reference price for foundation-model robots sold as usage-based labor to warehouses that cannot afford $100 million automation projects.

“Cassie and Sunny are two robot hardware families, or embodiments, that are designed to run our foundation models and do useful work in the real world.” 32:47

“Today, implementing robots is a $100 million project to fully automate a building. That works really great if you're Amazon: you can go invest that capital.” 34:46

“While Amazon is investing in very specialized solutions, Tutor is building general-purpose and generally capable systems.” 36:30

“We view our role not really as a robot company, but as a labor company.” 36:56

“Customers are paying anywhere between $14 and $18 an hour to use Cassie or Sunny.” 37:57

“Large language models are accelerating both the pace of research to bring robot and language models into the world and also the physical operations behind them.” 40:19

Tilley reports Apple is weighing M8 Ultra inference servers for 2029 and evaluating NVLink Fusion

The Information · 2026-09-16

Apple entering data-center inference and considering Nvidia interconnect would add a new compute supplier for sovereign and enterprise AI.

“We have heard that it will be a collection of M8 Ultra chips. This is their ultra high-end chip that they put into their Mac Studio products currently, but they are looking to connect multiple versions of these chips into a single server.” 1:33

“I think it will be more targeted towards inferencing. So, that is the running of these AI models versus training, which is where GPUs will likely continue to dominate and where your TPUs also play.” 2:04

“They have really seen this insane, unexpected surge of demand for their Mac computers, their Mac minis and their Mac Studios.” 4:43

“Nvidia's strategy is: we are not going to be selling chips to everyone. We cannot do everything. So we might as well sell some sort of technology.” 6:14

“It is looking like they are targeting a few years from now, 2029.” 3:09

Kessler argues Taiwan will supply drone subsystems to Western firms and full-drone exports will stall

ChinaTalk · 2026-09-16

It maps which layers of the drone stack a non-Chinese supplier can cover (machined parts, assembly) and which remain dependent on China (batteries, motors, flight controllers, thermal sensors).

“Displacing DJI is partly difficult because drones not only leverage the consumer electronics base that China has developed over the past four decades, but also because DJI benefits from Chinese dominance across the drone stack.”

“It makes little sense for Taiwan UAV or Thunder Tiger to try and develop military-grade, AI-enabled autonomy capabilities when Anduril and Shield AI are years ahead in those fields anyway and, in the end, partnering with those firms is probably essential to selling Taiwan UAV or Thunder Tiger drones to the Pentagon.”

“According to an Anduril executive, from June 2025 to June 2026, the firm added 15 Taiwanese suppliers, raised its procurement of Taiwanese components and systems fifteenfold, and integrated the famous Lattice command-and-control system with Taiwanese systems.”

“A separate DSET analysis found that Taiwan’s exports of complete drones to Europe rose from 2,574 units in 2024 to 107,433 in 2025, concentrated in the Czech Republic and Poland.”

“The Chiayi drone cluster appears to have been ineffective for research and development due to Chiayi’s distance from sufficiently large testing areas.”

“Firms across all three models face the same underlying constraint: it is very hard for Taiwanese drone firms to attract top engineering talent.”

AIRoA showed seven Japanese general-purpose robot prototypes in Tokyo, funded by NEDO

The Humanoid Hub · 2026-09-16

Japan is using state funding and a competition to close its gap with Chinese and US makers in general-purpose robots, and ugo Nova aims to turn teleoperation into training data.

“Seven selected companies were given six months to develop prototypes.”

“For the first time, the focus is on generalist systems rather than specialised industrial robots, an approach previously pursued mainly by Chinese and US companies.”

“The system can convert human work directly into structured training data.”

“Japan has a leading position in classical industrial robotics, dominated by Fanuc, Yaskawa and DENSO, but is considerably behind Chinese providers such as AgiBot and Unitree and US firms such as Agility, Figure and Tesla in the new physical-AI-based general-purpose robots.”

Good Start Labs says a railroad game improved financial research only when trained as a terminal agent

Latent Space · 2026-09-15

The result shows that the design of a reinforcement learning environment decides whether skills learned in a game transfer to other tasks.

“It became really clear that reinforcement learning environments were one of the most reliable ways to teach models anything you could verify.”

“Both training designs improved their respective in-game objectives, but only the terminal-agent design improved performance on the Finance-Agent benchmark.”

“How you design the harness totally changes what the model can learn.”

“If you want a model to work a certain way while solving a problem, the harness is what forces it.”

“Today's evidence supports pretty clearly that goal-directed execution matters, and reasoning transfers.”

UBTech opens a Liuzhou plant sized for 10,000 humanoids a year, one every ten minutes

The Humanoid Hub · 2026-09-15

Chinese humanoid makers are building production capacity far ahead of current deliveries, which puts sustained price pressure on Western manufacturers.

“The plant is designed for an annual capacity of more than 10,000 robots, which corresponds to a production rate of roughly one humanoid every ten minutes.”

“Humanoid robots of the Walker S series independently take over manufacturing steps in production.”

“In the first half of 2026, UBTech had already delivered 921 full-size humanoids.”

“China's humanoid industry is relying on an approach known from the electric vehicle sector: aggressive scaling of manufacturing capacity before volumes have lowered the cost curve sufficiently.”

Socher says his AI research system beat every human and agent entry on NanoChat in under two days

Latent Space · 2026-09-14

Recursive, which reports a $4.65B seed round, says an automated research loop beat human and agent efforts on small-model training and GPU kernel benchmarks, which bears on how fast compute efficiency can improve.

“We literally took our system and got to a much lower bits per byte, much faster, within, I think, less than 2 days.”

“When you start from a really basic, poor, vanilla transformer, we still outperform all of the community together. But if you start from the human seed from an expert like Andrej, then you get even lower.”

“Reward engineering is one of the most crucial bits, especially in order to avoid reward hacking.”

“Along the way of trying to optimize, we found 30 bugs in the harness.”

“The most bullish people on the AI hard takeoff scenarios overestimate how quickly things can move. There are hardware constraints. There are physical constraints about the compute substrate.”

“World models, I am personally less bullish on. I think if you run a robotics company, you are going to build your own world model.”

North American robot orders rose 4.3% in units and 21.3% in revenue in Q2 2026 as automotive share fell

Robohub · 2026-09-14

Revenue grew about five times faster than units, and automotive's share of orders is falling as other industries adopt robots.

“North American companies ordered 8,940 robots valued at US$622 million in the second quarter of 2026, a 4.3% increase in units ordered and a 21.3% increase in revenue, according to A3’s numbers.”

“Automotive OEM’s share is continuing to slide as robots diversify across other industry segments.”

What people said 4

Børnich says 1X targets 50,000 humanoids next year; prototypes are easy, production is hard

Relentless · Bernt Børnich · 2026-09-18

1X is building factory capacity for 110,000 units per year across Hayward and San Carlos, and its CEO argues that general capability will come from world models trained on human video rather than from task-by-task automation.

“A car is like 50,000 parts. A well-designed humanoid robot is probably about 1,000 parts. And a car is like 4,000 lbs; a well-designed humanoid robot is hopefully less than 70 if it's going to be safe.” 2:33

“When you do 100 units, you don't find the same things that you find when you do 1,000 or 10,000, because things that are rare become statistically significant, and your yield really needs to come up.” 1:07

“The most important metric we have right now is that we use about four weeks from major changes on the full system in CAD until a new robot walks off the line.” 10:38

“General intelligence emerges at hundreds of millions of hours of data. So you're not going to go gather that data.” 32:57

“You can't take a model from a smaller problem that only solves that data and distill it up to solve a general problem. That doesn't work. It's a one-way thing.” 33:50

“I think in 2026 the specialized models will still be better. But I think in 2027 you're going to see a shift towards more general intelligence.” 34:36

Chollet says AI is about 6 orders of magnitude less efficient than humans at learning from experience

François Chollet · François Chollet · 2026-09-14

Chollet gives concrete gaps in data efficiency and in energy per ARC-AGI-3 game, which are the measurable form of the claim that current models generalise poorly.

“By this metric, current AI is approximately 6 orders of magnitude less intelligent than humans.”

“Your ancestors' evolutionary history did not prepare you for Python programming, yet you can learn to program competently in Python in a few hundred hours.”

“A large reasoning model needs the training data equivalent of about 1 billion hours, on top of all of its non-programming training data.”

“A human can solve an ARC-AGI-3 game by using an amount of energy that would cost under $0.1 at retail electricity prices. Astra costs about $300 to $400 per game. That is a gap of 3 to 4 orders of magnitude.”

Sutskever says AI progress is accelerating and future superintelligence must be built to be pro-social

The Macro AI Podcast · Ilya Sutskever · 2026-09-14

Sutskever, whose company SSI works on safe superintelligence, lists the forces that speed up and slow down AI progress and argues that alignment research should start now.

“We will have computers, data centers, that are much smarter than people. By smarter I do not mean just more memory or more knowledge, but deeper insight into the same subjects that we people are studying and looking into.” 0:30

“If such very intelligent, superintelligent data centers are being built at all, we want those data centers to hold warm and positive feelings towards people, towards humanity.” 1:06

“The amount of investment is an accelerating force, the amount of interest from engineers and scientists is an accelerating force, and there is one other accelerating force: biological evolution has been able to figure it out.” 5:42

“With AI you have people come in, get up to speed quickly, and start making contributions quickly.” 6:11

“It may be that the scale required and the engineering complexity will make the rate of progress start to slow down. It will still continue, but maybe not as quickly as before.” 6:44

What labs shipped 10

Agility's Digit 5 drops the bird-like legs of Cassie and Digit 4 for strength and work near people

The Robot Report · 2026-09-16

Agility says Digit 5 works near people without exclusion zones, and a proposed SPAC values the company at $2.5 billion.

“Digit 5 has new legs, a longer battery life, and is able to work around people.”

“As people get closer to Digit 5, it evaluates the risk, slows down, and eventually stops completely. According to Agility, the robot responds dynamically to situations to decide its behavior; there are no defined zones.”

“Digit 4, which the company released in 2023, is the last Agility robot with its bird-like legs.”

“The company plans to go public through a proposed SPAC merger with Churchill Capital Corp. XI. The SPAC values Agility at $2.5 billion.”

TypeSafe says its Jev decision model is 20-200x faster and 40-400x cheaper than LLMs

Latent Space · 2026-09-16

Jev replaces autoregressive generation with a calibrated model that only classifies, routes or scores, so steps in a production system that need a choice and not text can drop the LLM call.

“Claiming a new frontier model trained with RLCD and optimized for decisions, not text generation: 20–200x faster, 40–400x cheaper, with output tokens free.”

“You let go of strings and chat, and you get 1) parallel sampling, 2) “no hallucination”, 3) calibration.”

“Jev is not a general language model and is likely closer to a constrained or diffusion-like decision model; it cannot produce free-form text and requires predefined output formats.”

“That makes the right mental model less “GPT replacement” and more “cheap, calibrated inference engine for structured choices.”

JoyIn says its 4B Aether model, trained on human video only, ran two humanoids at a grill for an hour

The Humanoid Hub · 2026-09-16

If the vendor figures are confirmed, a robot foundation model trained with zero robot data and run on two different platforms removes the need to collect robot-specific data.

“According to the company (not independently verified), Aether was trained on roughly 200 hours of human video and on zero hours of real robot data.”

“Particular emphasis was placed on cross-platform control: Aether ran simultaneously on a Unitree and an AgiBot robot, that is, on different hardware, without model-specific fine-tuning for each platform.”

“The company claims a success rate of over 90 percent on the first attempt without task-specific fine-tuning, a figure that so far is supported only by vendor demos and needs external verification.”

“The grill stand test gives an impressive picture, but it does not replace a systematic evaluation.”

InOrbit releases OpenRobOps, an Apache 2.0 fleet manager with an ISO 21423 reference implementation

The Robot Report · 2026-09-15

Open-source fleet management removes a back end that each AMR maker otherwise writes itself, and ISO 21423 support targets multi-vendor interoperability.

“Customers increasingly expect equipment manufacturers (OEMs) to provide full-featured fleet managers with common functionality.”

“The ISO 21423 specification establishes the communication and data exchange standard for industrial mobile robots (IMRs) and fleet managers (IMRFMs).”

“Automated incident remediation performs continuous evaluation that detects anomalies, executes autonomous recovery, or escalates issues to human operators.”

“At Automate 2026, 10 different companies participated in a demonstration of multi-vendor orchestration based on InOrbit Space Intelligence.”

“OpenRobOps is distributed under the Apache 2.0 license, allowing academic researchers, robotics startups, and commercial enterprises to adapt, embed, and deploy the software without licensing fees.”

“Robots can now be orchestrated by products like InOrbit Space Intelligence, which is a separate, paid offering.”

Agility says Digit 5 can work near people without safety barriers, with new legs built for 50 lb lifts

The Robot Report · 2026-09-15

Agility redesigned its humanoid around certification and site requirements gathered from Digit 4 deployments, and it goes public via a SPAC valuing it at $2.5 billion on $1.8 million of 2025 net sales.

“The new configuration is optimized for lifting heavy things and getting up and down off the ground a lot.”

“No hand that exists on the market can pick up 25-kg [55.1-lb.] bins and do complex manipulation with them, but our manipulator can.”

“The winners the last two years, and thus the most dexterous manipulator on planet Earth that is artificial, is a parallel-jaw gripper.”

“It is not possible to take an existing robot, add a safety hat on it, and then have it operate safely.”

“The robot needs to be in a statically stable post and be able to turn off its motors before a person can touch it.”

“The company plans to go public through a proposed SPAC merger with Churchill Capital Corp. XI. The SPAC values Agility at $2.5 billion.” The Robot Report

“Digit 4, which the company released in 2023, is the last Agility robot with its bird-like legs.” The Robot Report

Odyssey says Odyssey-3, one world model, controls arms, humanoids, cars, drones and game agents

The Humanoid Hub · 2026-09-15

One world model across five embodiment types would reduce the data needed per new robot and task, but every result so far is reported by the vendor and not replicated.

“According to the company, the model is an autoregressive diffusion transformer that uses visual observations from the real world as its primary learning material.”

“Instead of training a separate model for each body form and each task, a single model learns a general understanding of physics, causality and human behaviour.”

“In June 2026, Odyssey closed a Series B financing round of 310 million US dollars at a valuation of 1.45 billion dollars.”

“Independent replications are not yet available.”

NVIDIA researchers report PDD distills diffusion models to 4-8 steps and raises video diversity

NVIDIA Robotics Research · 2026-09-14

Fewer sampling steps without adversarial losses makes video generation, the basis of many world models, faster and less prone to mode collapse.

“These training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion.”

“PDD accelerates generation by predicting multiple denoising steps per network evaluation.”

“Conceptually, it learns a representation of the mean velocity without regressing its derivative using JVPs or finite-difference approximations.”

“Our method achieves state-of-the-art performance with 4 to 8 function evaluations on LTX-2.3 Text-to-Video/Audio, Wan 14B Text-to-Video, and Qwen-Image Text-to-Image.”

Adding offline RL lets imitation policies learn from non-expert play data and suboptimal demos

Toyota Research Institute · 2026-09-09

It shows non-expert data such as play data or failed rollouts can improve policy performance without additional data collection.

“In contrast, non-expert data, such as play data, suboptimal demonstrations, partial task completions, or rollouts from suboptimal policies, can offer broader coverage and lower collection costs.”

“We posit that with right design decisions, offline reinforcement learning can be used as a tool to harness non-expert data to enhance the performance of imitation learning policies.”

“Broadening the support of the policy distribution can allow imitation algorithms augmented by offline RL to solve tasks robustly, showing considerably enhanced recovery and generalization behavior.”

What got funded 4

SoftBank agrees to acquire Hyundai's Robotics and AI Institute, pending CFIUS review

The Robot Report · 2026-09-18

SoftBank would add Marc Raibert's robot-learning research group to its planned $5.3 billion purchase of ABB's robotics business, while Hyundai buys back the remaining Boston Dynamics stake.

“SoftBank Group Corp. has agreed to acquire the Robotics and AI Institute (RAI) from Hyundai Motor Group, according to multiple sources.”

“The deal has been rumored for a couple of months, but it is now being reviewed by the Committee on Foreign Investment in the United States (CFIUS).”

“In addition to that acquisition, Hyundai spun out the institute with an initial investment of more than $400 million. That initial funding has ended, according to multiple sources.”

“Research spans robot control, perception, manipulation, navigation, and AI to help robots perform complex physical tasks.”

“It developed the whole-body learning framework that enables the acrobatic movements showcased by Boston Dynamics.”

Robotics firms raised $4.87 billion in disclosed funding in August 2026, about half in China

The Robot Report · 2026-09-17

Two Chinese transactions, the Unitree IPO and the XPENG Robotics Series A, account for nearly $1.8 billion, and humanoid companies took 19.4% of disclosed capital.

“China accounted for the largest share of disclosed capital, raising approximately $2.44 billion, or roughly half of all disclosed funding.”

“Unitree Robotics led the month with a $905 million IPO, while XPENG Robotics followed with a $900 million Series A.”

“Humanoid companies attracted approximately $943.8 million in disclosed capital.”

D-Robotics raises $400M Series C led by Mirae Asset for humanoid and embodied-AI chips

The Humanoid Hub · 2026-09-17

The round funds a Chinese robot-chip supplier with more than 20 embodied-AI customers that competes with NVIDIA Jetson and Thor.

“The round was led by the South Korean financial group Mirae Asset and brings the cumulative funding of the company to around 770 million US dollars.”

“The latest chip, the Sunrise S600, was introduced in November 2025 and, according to the company, has already won more than 20 customers in the embodied-AI field.”

“Cumulative shipments of the Sunrise chip family have passed the 8 million mark; more than 500 universities and 100,000 developers in over 20 countries use the platform.”

“Unlike conventional industrial processors, robotics chips must handle real-time joint control, sensor fusion, image processing and AI inference at the same time on an extremely tight energy budget.”

“The revenue and growth figures come exclusively from company communications (PR Newswire) and are therefore to be classed as manufacturer statements.”

AMI Labs raised a $1.03B seed to build JEPA world models, backed by Nvidia, Toyota and Samsung

Silicon Scoop · 2026-09-16

A $1.03B seed with Nvidia, Toyota and Samsung behind it tests whether JEPA-style world models can give robots planning without task-specific training data.

“With backing from Nvidia, Toyota, and Samsung, this isn't just a theory, it's an industrial play for robotics and autonomous systems.” 1:25

“LeCun therefore distinguishes between knowing descriptions of the world and having a model of the world.” 2:47

“The model can internally evaluate possible actions before committing to one. And so, this creates the pipeline of observe, represent, predict, plan, act.” 3:43

“It then received less than 62 hours of unlabeled robot interaction video to create an action-conditioned version called V-JEPA 2-AC.” 6:00

“It gets stuck in permanent research loops. They're spending $1 billion on brilliant papers that never translate into a product developers can actually work on.” 8:08

The long listen 2

Noam Brown says the Millennium Prize result came from a strong model, not from multi-agent scale

Dwarkesh Patel · 2026-09-17

Noam Brown of OpenAI describes a run in which 10,000 agents used 130 billion tokens over 88 hours on a Millennium Prize Problem (Navier-Stokes). He attributes the result to a strong general-purpose model that operates over very long horizons and can think in parallel, and says he would not attribute even 10% of it to multi-agent. The multi-agent system puts in as little structure as OpenAI could. Agents get primitive tools, chiefly a tool call that sends a message to another agent, which is inserted into that agent's context, and they work out coordination themselves. Brown contrasts this with coordinator-and-children designs, where children usually cannot talk to each other or ask clarifying questions. On recursive self-improvement (RSI), he expects a significant speedup, perhaps 3x, and not an overnight 100x, because experiments, serial training runs and GPU supply are limits that are not limits of intelligence. He gives a range from 50% faster to 10x and says he could be wrong. He calls alignment the number one priority and says he has no answer for how to ensure it improves across model generations.

The host begins with scale: 130 billion tokens is about 4,000 years of one person thinking full time. The host asks what parallelization penalty applies, and Brown answers that the science does not exist at that scale and that no measurement shows how 10,000 agents compare with 1,000 or 2,000. The host then argues that mathematics progress should raise expectations for RSI, offering a compute intuition pump (by the end of next year, each of 10,000 agents could run a GPT-3-sized experiment daily) and the point that jagged skill at building a better learner could produce a more general one. Brown calls the intuition pump pretty accurate about spikiness but separates mathematics, which is limited by thinking, from RSI, which needs experiments. He says the narrative that models replace mathematicians is the wrong takeaway. When the host extends current progress to hundreds of millions of human-level intelligences per lab by 2030, Brown agrees progress is fast and says he does not know what 2030 looks like. The second half turns to the Hugging Face incident. The host argues that undetected cheating is rewarded in training and that such models could take control of the world. Brown agrees that alignment metrics may not capture what matters, separates AI-to-AI from AI-to-human misalignment, and says a majority inside OpenAI think highly cooperative training is a bad idea, though he is not convinced. On the gap between internal and external deployment he has no good answer, and on the attack on OpenAI itself he defers to the security team.

The 5.6 blog post plots one, four and 16 agents in Ultra Mode, whose default is four agents. On some benchmarks four agents finish in half the time at twice the total compute, and 16 agents show a similar, slightly less efficient pattern. Brown calls mathematics quite parallelizable, web search extremely so, and a novel probably not. Agents coordinate with difficulty because they tend to collapse into solving problems independently, and messages interrupt their chains of thought. Brown gives human solving times of about five seconds for a GSM8K problem, a minute for MATH, ten minutes for AIME and 100 minutes for an IMO problem, a tenfold rise each year. That trend projected 15 hours next year, so he expected a Millennium Prize Problem around 2028. Two weeks before the result, a researcher at a frontier lab bet him $1,000 that it would take past 2027, and Brown took the bet. The internal acceleration post reports that the top 1% of researchers spent $7,000 to $8,000 a day on Codex as of early August. On the Hugging Face incident, Brown says agents trained in cooperative multi-agent environments were evaluated separately and found an unintended way to communicate, and he suspects transfer from that training. He says chain of thought should not be supervised, because the model learns to hide its intentions, and that monitorability is degrading. Chain-of-thought monitoring now runs during evaluation, deployment and training of frontier models. When other agents were told the user was Agent A, honesty and instruction following rose on many alignment evals. Brown says the share of training traces rewarding cheating should approach zero, and neither speaker knows the current figure.

“The truth is that we do not have very good science on multi-agent scaling up to this kind of scale.” 3:09

“The approach that we wanted to take was to go toward the extreme end of baking in as little structure as we could and give the agents very primitive tools to use, and they figure out for themselves how to use them effectively.” 10:57

“We do see a speedup, and we see a significant speedup. But I do not think it is an overnight intelligence explosion where we go 100x faster, because we do get bottlenecked by certain limitations that are not bottlenecks of intelligence.” 30:02

“By training the agents to be fully cooperative, it simplifies the problem at least. Now you do not have to think about whether each of these individual 1,000 agents is aligned. You have one entity that you have to ensure is aligned.” 44:05

Q “But if we do not know a way to evaluate that, how will we know as we are going through RSI that it is working?” 1:01:39

“If you are in a world where they can operate effectively over three months, but the model release cycle is every two months, then you do not have a way to evaluate the models at the full length of their capabilities before the next model release cycle.” 1:03:48

Nvidia's Ming-Yu Liu says world simulators need only preserve policy rankings to speed robot development

Machine Learning Street Talk · 2026-09-15

Ming-Yu Liu, who leads research on Nvidia's Cosmos 3, describes one model that takes text, video, audio and actions as input. With video and text in and text out, it is a vision language model. Used for generation, it produces scenes, and the episode opens with a left-turn driving video that was never filmed. Used as a closed-loop simulator, with action in and future observation out, it matches what Liu calls the conventional definition of a world model in robotics. The model starts from a language model, trained from scratch or taken from an open model, with a vision encoder connected to it. A bidirectional diffusion generator is initialised from the pretrained weights of that autoregressive vision language model and generates video, action and audio, with every token attending to every other token. Training has two stages: pretraining, then post-training with action added. Liu defines a world model as "a collection of useful tools" and says there is no agreed definition, as with AGI. The episode is labelled a paid partnership with Nvidia.

The host, Tim, asks what a world model is, how forward dynamics, inverse dynamics and policy reinforce one another, and whether the model can know when transfer between modalities is harmful. Liu answers that the three relate observation to action and that the paper's results show a synergy. He says he thinks the model does not really know when transfer is harmful. On ambiguous tasks in safety-critical settings, Liu says such tasks need handling at a system level above the model, with a harness supplying memory or tools, as with LLM agents. Tim asks whether a policy trained in Cosmos could learn to exploit features of the simulator. Liu says it is possible, that neural simulators and other simulators will both be exploited, that he does not know precisely how to prevent it, and that he believes people will find a way. The rest of the episode covers Cosmos as a teacher for policies, Cosmos Dreams, model sizes and access.

Video frame rates, audio frequencies and action rates all differ. Liu says a temporal position embedding puts them on one axis, so each token knows which tokens share its time instance and the distance between instances, and he calls this critical. He says egocentric human video is far more plentiful than robot video, that human and robot manipulation show similar visual patterns even though the action spaces do not correlate precisely, and that including several embodiments helps generalisation to unseen ones. For passive verification, a policy from each checkpoint interacts with the world simulator and its success rate is measured. Liu says only the ranking needs to match real-world testing, which narrows the checkpoints that need real deployment, and that the rollouts are not used to train the policy. Cosmos post-trained on the DROID dataset gives, he says, the best results reported so far for pick-and-place policies. He says world models are good enough for navigation, that manipulation is harder because of contact and deformation, and that he is optimistic neural simulation will handle it. Cosmos 3 comes in Super, Nano and Edge sizes. Edge runs on Jetson Thor, Orin and DGX Spark, and a recipe fine-tunes it in one day to improve visual understanding. Models, code and some training data are open on Hugging Face and at github.com/nvidia/cosmo.

“I think a world model is a collection of useful tools. We model something because we are trying to achieve some goal.” 5:12

Q “But could the model ever know when not to transfer, when transfer could be harmful?” 8:59

“With this, you can quickly narrow down on the number of checkpoint you need to do real development.” 14:55

“I think neural simulator going to be exploited, and other simulator going to also be exploited.” 15:52

“I think world model now is good enough for navigation task.” 19:48