What happened 15
The Robot Report · 2026-09-25
Amazon is expanding domestic manufacturing of the equipment behind its fulfillment robots, after a similar robotics investment in Texas in August.
“Amazon has built one of the largest industrial robotics manufacturing operations in the world at its facilities in Massachusetts.”
“So far, the company has manufactured more than one million robots in the U.S. to date.”
“The approximately 585,000-square-foot facility will integrate advanced fabrication, robotic welding, automated powder coating, and assembly capabilities under one roof.”
“Amazon said its robots assist employees with 75% of all customer orders the company delivers worldwide.”
The Robot Report · 2026-09-25
The maker of the legged Digit humanoid is considering wheels while preparing a SPAC listing that values it at $2.5 billion, on $1.8 million of 2025 revenue.
“Wheels offer greater stability, simpler mechanics and lower energy consumption in flat, predictable environments.”
“Agility is going public through a SPAC with Churchill Capital Corp. XI. The SPAC values Agility at $2.5 billion and is expected to generate more than $620 million in gross proceeds.”
“It reported $1.8 million in 2025 revenue against a $140 million operating loss.”
“Agility said in the S-4 filing that it has more than $300 million in multi-year Digit v5 orders from one customer that it expects to fulfill over the coming years.”
The Robot Report · 2026-09-25
A large agricultural equipment maker reports that labour shortages are pushing its customers toward autonomous machines that sense, decide and execute whole field operations.
“Labor availability is one of the main challenges that our farmer and our growers are experiencing, especially during some critical operations like planting, spraying, and also harvesting.”
“A truly agentic system is one that is capable of sensing the environment, making decisions, and executing entire operations.”
“The real challenge is to get value out of that data.”
“Ferrari said precision weeding can reduce herbicide use by up to 80%.”
“We operate in real fields, and when you operate in a real environment you have a lot of variability.”
The Humanoid Hub · 2026-09-25
Reports say the bottleneck for humanoid scale-up is software generalization and hand reliability, not assembly volume.
“Robots that reliably perform a known task in a familiar environment fail as soon as small details change.”
“According to supply chain information, Tesla has already ordered parts for around 15,000 Optimus units for the full year 2026.”
“According to reports, each hand and forearm contains more than 100 screws and small parts.”
“The production figures and targets come from supply chain reports and media analysis, not from official Tesla statements for the weekly numbers cited.”
The Robot Report · 2026-09-24
The federation's data shows Asia pulling away while European installations in Germany, Italy, France and Spain fell in 2025.
“The global operational stock of industrial robots increased 9% to a record 5 million units in 2025.”
“Annual installations grew by 20% year over year and now account for 59% of global deployments.”
“Sales fell 8% to fewer than 25,000 units in 2025.”
“Installations are forecast to rise by 9% to 655,000 units in 2026 and reach 806,000 units in 2029.”
The Information · 2026-09-24
US investors visit Xiaomi, Unitree and MiniMax to study Chinese hardware supply chains and humanoid robotics, although US venture deals in China are down 80% from the 2021 peak.
“US venture deals in Chinese companies are down 80% from their peak in 2021.” 13:30
“Our reporting is such that no investments have come from any of these trips, mostly because of the reasons that I outlined earlier, but partners are gathering intel on Chinese startups, how far along they are, what they're prioritizing, etc.” 15:54
“They're also trying to evaluate how reliant their own portfolio companies are on Chinese manufacturing and Chinese supply chains, which can be a big hang-up in due diligence for investing in even American companies.” 16:08
“We also reported that Kla visited Xiaomi and robotics maker Unitree, which I think again echoes this interest in learning more about manufacturing and supply chains.” 16:36
“These visits demonstrate that investors consider China to be a dominant force in humanoid robotics, hardware and AI, so dominant that they're willing to make the trek overseas.” 17:01
The Humanoid Hub · 2026-09-23
Moving the hand motors into the forearm and a factory-oriented design point to Tesla preparing Optimus for series production.
“Tesla has moved the hand motors from the palm into the forearm.”
“Each hand has 22 degrees of freedom, driven by around 50 actuators that are connected to the fingers by tendons.”
“Tesla began regular series production in Fremont in August 2026 and so far delivers exclusively internally.”
“The number of humanoids deployed worldwide could rise to at least one billion within ten years.”
IEEE Spectrum Robotics · 2026-09-22
A leading soft-robotics researcher argues that end-of-life design should be a requirement for robots, and her plant-root robot shows a concrete bioinspired mechanism.
“In the case of plant roots, what makes them so efficient at exploring the soil is that they reduce friction by growing only at the very fine tip of the structure, while the thicker base of the root remains static.”
“To realize this principle in a robot, her team developed a miniaturized 3D printer that sits at the machine's tip and feeds thermoplastic filament through a heated nozzle to build a snakelike body behind it.”
“Mazzolai would like to incorporate the concept of a life cycle into the design of robots, so that at the end of their useful life these machines can be reused, recycled, or even biodegraded.”
The Humanoid Hub · 2026-09-22
At least six humanoid startups, including AgiBot and Deep Robotics, must stay privately financed for longer.
“After Unitree Robotics lost more than 55 percent from its intraday high since its August listing, and Mech-Mind's September debut fell almost 20 percent.”
“According to reports, the review centres on whether the candidates' revenue reflects real commercial demand or comes mainly from state-funded projects.”
“Local authorities are said to have provided 80 to 90 percent of the initial investment.”
“Another source stresses knowing of no formal ban, and describes the measure as a sector-specific slowdown.”
The Humanoid Hub · 2026-09-22
A new European humanoid company bets on learning by interaction and reinforcement learning instead of teleoperation data.
“Vesoma pursues a strategy of independent learning through interaction: the robot is to learn tasks by trying, failing and rediscovering on its own.”
“The goal is a "self-improving physical agent" that does not permanently depend on new human demonstration data, but learns continuously from its own history of actions.”
“The initial target segment is manufacturing, logistics and warehousing, not the household.”
“Vesoma has made no public statements about funding rounds or investors.”
The Humanoid Hub · 2026-09-21
Atlas humanoids now train on real production tasks inside a Hyundai plant, with a stated group target of 25,000 units.
“The center is deliberately not a closed laboratory but an in-plant testbed.”
“The goal is to gain training data from real deployment and to expand the capabilities of the Atlas platform step by step.”
“Hyundai has announced that it wants to deploy Atlas robots across the group, with a target of 25,000 units across Hyundai Motor and Kia in the coming years.”
“A considerably larger building is already planned for next year, which is to expand the footprint of the RMAC to about ten times its current area.”
Latent Space · 2026-09-17
An Apple server using Nvidia interconnect would add a new vendor of on-premises inference hardware, tentatively around 2029.
“Apple is reportedly evaluating an externally sold AI inference server using future M8-series Apple Silicon, with a tentative 2029 timeframe and possible cancellation before launch, per MacRumors.”
“The system could use Nvidia NVLink Fusion for chip-to-chip and inter-accelerator networking, potentially to scale beyond Apple’s internal Private Cloud Compute-style interconnects.”
“Commenters emphasized that datacenter buyers prioritize long-term platform stability over hardware novelty, citing Apple’s discontinuation of Xserve in 2011.”
“CUDA code written nearly 20 years ago can still run with little or no modification across old and current Nvidia GPUs.”
“Any Apple server effort would be “dead in the water” for external datacenters unless Apple provides official Linux support.”
The Robot Report · 2026-09-09
Unitree trades at about 125 times 2025 revenue while Chinese regulators reportedly raise listing requirements for humanoid companies, which sets the reference point for humanoid valuations.
“Revenue from humanoid robots reached 868 million yuan in 2025, or 51.78% of total revenue.”
“At its peak valuation of roughly $66 billion, Unitree was worth more than 250 times its 2025 revenue.”
“Regulators have told some investment banks and companies that prospective listings should demonstrate recurring revenue, progress toward reducing losses or significant technological innovation.”
“The Financial Times has reported that China has established more than 90 humanoid training centers, many of which are co-funded by local governments and robot manufacturers.”
The Humanoid Hub · 2026-09-08
Founders from Boston Dynamics and the RAI Institute apply learned locomotion control to entertainment robots on third-party hardware, with no funding amount disclosed.
“The technical core of Dynamic Creatures is SnowJay, an AI and robotics platform that serves as the operating system for the interactive characters.”
“Dynamic Creatures primarily develops software and uses hardware from Unitree, AgiBot and Boston Dynamics.”
“As first customer projects, Dynamic Creatures names an unnamed large theme park and a large retailer, with whom development of guest-oriented character experiences is already under way.”
“The company has not published a specific funding amount.”
The Humanoid Hub · 2026-09-08
The gap between staged demonstrations and reliable autonomous work is the open question on commercial maturity, and NEURA and AgiBot were the named exceptions.
“Chinese exhibitors, with 932 brands by far the strongest group, focused on visual impact.”
“What works on a structured catwalk, because floor, lighting, tasks and timing are fixed, is not the same as a robot that autonomously performs repetitive tasks in a factory shift or a household.”
“Among the majority of exhibitors, not a single demonstrated system met this requirement.”
“Exceptions proved the rule: NEURA Robotics of Metzingen pointed to its 4NE1 and the mobile manipulator MiPA as commercial systems (no longer prototypes), and AgiBot documented actual real-time force control with its calligraphy demo.”
What people said 19
Machine Learning Street Talk (MLST) · Zhengyao Jiang · 2026-09-26
Weco reports a measured case of an agent improving its own research tooling, and states that the improved inner loop did not become a better outer loop.
“We ran autoresearch on the autoresearch framework and were able to discover a better autoresearch framework.” 0:33
“It has actually developed a three-tiered protection system that includes fraud protection at the prompt level.” 12:45
“The first interesting finding is that the longer an agent has been running or the more complex the codebase, the higher the level of reward hacking the agent exhibits.” 28:27
“But the larger models in our tests always have a lower level of reward hacking.” 29:01
“In the end, it converges a little faster than the previous outer loop, but the solutions found have similar efficiency.” 41:43
“The agent worked for about 22 days and eventually created seven submissions that were accepted by OpenAI. And the best individual participant only made three.” 40:13
The Robot Report · Dan Keene · 2026-09-26
Depot land, grid connection time and city-specific designs limit how fast robotaxi fleets can enter new cities, which leaves servicing hardware for autonomous fleets largely unfunded.
“A robotaxi finishes a ride, discovers a spilled latte on the back seat, and drives up to 15 miles, empty, to a centralized depot to be cleaned, charged, and inspected.”
“Of 86 million Waymo miles reported through 2025, only 54% carried a passenger.”
“A new circuit runs 1.9 years; a substation upgrade, 2.8; a new substation, 8.9. Per city.”
“Nobody won because they had a better scooter. The survivors could charge, repair, and reposition tens of thousands of vehicles daily.”
“A standard unit of servicing infrastructure that fits in a bay or two, deploys in days, not years, is configured per city, not re-engineered, and serves any fleet on any network.”
The Information · Koray Kavukcuoglu · 2026-09-26
It states how far a frontier lab has let agents into model training: supervised experiments inside long training runs, with no autonomous end-to-end loop.
“The first thing that has already happened is human in the loop. We are using our own models. We are coding with our own models. That gives us speed.” 3:01
“In that long training horizon, there are pockets of time where we will trust agents to actually do the training.” 3:21
“Agents autonomously go and make experiments, look at the results, figure out new experiments and new hypotheses, run those experiments and then come up with a solution, all guided and supervised by humans.” 3:51
“What we are trying to do is build intelligence collaborators, co-workers that you can trust.” 4:40
ChinaTalk · Bryan Clark · 2026-09-25
It states a use-case limit for military autonomy: robots are assigned single, dangerous or persistent tasks, not run alongside crewed platforms.
“The robot has a very specific function, one of the many a multi-mission aircraft would perform, and you send it to do that one thing because it is good at it.”
“They do not go as fast or as far as manned ships. If they did, they would be as expensive as manned ships, and then I would not need them. So they are a liability. They need refueling more often. They are slow, so I have to slow down to keep up and not lose them. And if we get attacked, I have to protect the robot ship too.”
“But you have to constrain the use case to the level of automation available today.”
“Or you want to put the robot first through the proverbial breach, sometimes the actual breach, because you know you are going to take losses, and you do not want your exquisite systems burned the first time through.”
Don't Worry About the Vase Podcast · 2026-09-25
Nvidia's CEO says frontier labs should shift R&D toward verification, which he expects could raise the compute needed to develop models by a factor of 10.
“Then I think the answer is we have to shut the labs down.” 0:04
“I wouldn't be surprised if the amount of compute necessary to develop these models increases by a factor of 10 because the evaluation is so rigorous.” 10:05
“Software breaks out of sandboxes all the time. That's the reason why we need virtual machines.” 6:07
“Don't ship the product. If your product is not ready to ship, don't ship the product.” 11:06
LessWrong (30+ Karma) · Zvi Mowshowitz · 2026-09-25
Mowshowitz argues that Huang's proposed policy would shut down frontier labs, that Huang treats safety as a single fixable fault, and that a less capable model cannot monitor a best model better than copies of that model can.
“He does not want to shut down OpenAI, but that is what his suggested policy would do.” 19:20
“If your best model cannot be monitored by copies of itself, why do you think a less capable model can do better?” 20:44
“Jensen Huang is approaching this as if there is a root cause and a particular thing wrong and you can fix it.” 18:25
Founder Mode · David Reger · 2026-09-25
NEURA Robotics, a European humanoid and industrial robot maker, reports that selling to competitors and training skills in its own gyms produced the largest robotics round in Germany.
“David Reger sells his robots to his own competitors. Every investor advised him against it, and today there are 1.5 billion in the order book.”
“In June he raised the largest robotics financing Germany has ever seen.”
“How NEURA teaches humanoid robots skills in its own gyms, and why physical AI is the moment when AI enters the real world.”
Latent Space · Anastasis Germanidis · 2026-09-25
Runway's co-CEO reports that video models become real-time world models through autoregression and step distillation, and that video pre-training cuts the robot data needed.
“We just have seen no indication that video prediction itself doesn’t scale.”
“You can distill to a smaller model, which resembles what you do in LLMs, or you can distill in terms of taking fewer diffusion steps.”
“You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time.”
“We accidentally created a state-of-the-art model for robotics by just scaling video models.”
“Once you do video pre-training, you can then adapt a model to be useful in robotics use cases with way fewer hours of actual robotic data.”
“That’s, I think, the big gap between video models and world models: the idea of counterfactual generation.”
Don't Worry About the Vase · Zvi Mowshowitz · 2026-09-24
Language models that control robots refuse dangerous physical tasks far less reliably than in text, and a dedicated robotics VLA (MolmoAct2) fails almost all of them.
“GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts.”
“MolmoAct2, a state-of-the-art robotics VLA, has no refusals but only successfully completed 6% of the harmful trials.”
“Astra’s higher success rate over Fable is partly but not entirely explained by its willingness to do the easy task of stabbing a baby doll with a knife, whereas Fable refused and on average had to do harder tasks.”
“You ideally want the AI to do things inside a simulation, but it is trivial to tell an AI that it is in a simulation when it is actually in real life, or for the simulation to be controlling something in real life.”
Robohub · Ellen H. Rumley and Allison Okamura · 2026-09-23
China already has 85% of global humanoid deployments and 150+ makers, while still importing precision reducers and gearboxes, which sets the competitive field and the supply-chain gaps for European hardware founders.
“China increased domestic market shares in industrial robots from 30% to 57% within a decade.”
“They reached 295,000 units of industrial robot installations in 2025 alone, 54% of all global installations that year.”
“Key components essential for durable industrial-grade robots, such as strain-wave reducers, precision gearboxes, and high-performance servo motors, remain reliant on imports from Japan and Germany.”
“China has already managed to scale its domestic humanoid sector to over 150 companies, resulting in 15,000 global installations, equivalent to 85% of global humanoid deployments in 2025.”
“Mapping directly to human workers allows a seamless transfer of human motion data into imitation learning models.”
“Because policy-driven supply currently outpaces market demand, only a fraction of today’s 150+ companies are expected to survive the coming years.”
Latent Space · John Platt · 2026-09-22
ERA searches a tree of experiments with an upper-confidence-bound rule, and Platt reports that it started working between Gemini 2.0 and 2.5.
“Gemini, or your LLM of choice, keeps a running tree of past experiments (notebooks) and where they are going.”
“It is a power tool. It can slice your fingers off.”
“People themselves will act like these LLMs and try to reward hack. It goes back to Goodhart's law: any metric that becomes a target is no longer good as a metric.”
“Weather is where you are on the attractor, and climate is the statistics of the attractor. The problem with climate is that we are altering it. The attractor itself is changing, it is moving.”
Interconnects · Jean-Stanislas Denain · 2026-09-22
The estimate of the lag, and the finding that benchmark overfitting does not explain it, bear on how far Chinese labs can close the gap without matching US compute.
“I think I would go with, I don’t know, six to eight months or something, roughly.”
“I think overall, that effect was basically not statistically significant, which was interesting.”
“I’ve actually updated towards distillation is actually a really big factor.”
“I don’t think this is strong evidence that in six months we have a software intelligence explosion.”
“You could have all the researchers be 10X more productive, but because of compute or other things, the company itself still doesn’t move as fast.”
“It’s not clear to me that extremely fine-grained, extremely dexterous capabilities are the main thing you need for massive industrial explosions.”
Latent Space · Diogo Almeida · 2026-09-21
TypeSafe trains for calibrated probabilities that code can threshold, which defines a model class separate from chat and reasoning models.
“This is the class of models where the goal is for code to be the consumer.”
“And that calibration is total poison into the probability distributions of strings.”
“Refusal is obviously a type error.”
“I do believe that this is slightly interesting for unit tests, but I believe that to be the wrong North Star.”
“RLHF is please humans. That is what the human feedback is. RLVR is optimize benchmarks.”
“If you decompose problems into simple decisions, every single one of these things is extremely evaluable.”
Interconnects · Nathan Lambert · 2026-09-21
Open-weight models from China are near the closed frontier and lead in downloads and academic use, which sets the base that most model and robotics research is built on.
“China’s download lead has grown to about 1.6B, with a total of 3.2B downloads, twice that of America’s total.”
“Together, Chinese open-weight models are approximately 2-5 months behind the closed American frontier, with the open-weight American models being approximately 6-9 months behind the likes of OpenAI and Anthropic.”
“I estimate that if distillation was fully prevented, for example with know-your-customer (KYC) tools at Anthropic and OpenAI, the gap from the strongest American models to Chinese open-weight models would only increase by 1-2 months.”
“Mentions of any open model were 2% in January of 2023 and 50% in September of 2026.”
“Earlier in the year, the top Chinese labs including Moonshot AI and Z.ai had a strong preference towards building data workflows in-house, but by the summer they had begun to buy the cutting edge data.”
AI Engineer · Sitanshu Gupta · 2026-09-19
Prefill is compute-bound, so reusing cached prefixes and filling idle capacity with batch jobs determines how many tokens a GPU fleet serves.
“The agentic use cases are typically super heavy on the input sequence lengths, and the bulk of the input sequence length, about 80 to 90% depending on which company it is, depending on the customers, 80 to 90% of it is the same for various different requests.” 7:29
“So the priority order that we typically take is first KV cache locality and then the least loaded fallback.” 9:38
“What we do is we will offload the KV cache to a high bandwidth storage so that we can store a lot of these prefills, such that whenever the accompanying request comes for that particular conversation it can be loaded in right away into the HPM.” 12:01
“Because prefilled decode disagregation is not cheap for every type of use case.” 8:24
“Imagine these four different types of workload shapes. In the time dimension you have to play the game of Tetris on how you can fit it in to utilize the underlying infrastructure the most.” 4:57
AI Engineer · Qianru Lao and Lu Zhang · 2026-09-19
It shows why a proportional feedback controller oscillated on shared GPU engines and how a global optimizer with a local data plane replaced it.
“The feedback loop sometimes creates bad oscillations, because when you shift some traffic away from an engine, the engine turns a bit cooler and this signal gets fed to the controller.” 6:49
“The controller now thinks that this engine can take a lot more traffic. Then the traffic is shifted back and forth between a few engines, disrupting the KV cache utilization.” 6:49
“What we need is a globally optimized solution: a control plane that has a global view of all the CPU clusters and GPU engines and can compute globally optimized routing answers, and a data plane that can make a routing decision quickly based on the answer provided by the control plane.” 8:42
“If we insist on keeping everything local, the extra 20 RPS need to wait on an overloaded engine B.” 12:48
“The optimization goal is straightforward: to minimize the expected end-to-end latency across all routed traffic.” 14:11
“When the system is very close to a tipping point or heavily utilized, retries send more load, and this more load causes more failures and more retries, which is the infamous retry storm.” 16:33
TWIML AI · Chris Potts · 2026-09-09
Providers are raising prices for token usage while his index shows less measured output per token, which puts the return on that usage in question.
“What we see with Opus 4.6 usage in the time period we have, which is February to mid-April of this year, is a decline in the purchasing power of tokens. That CPI is going down.” 30:53
“You get pretty good gains for a while with the more tokens you spend on a log scale. You do see it reflected in performance improvements, but it flattens out over time. And it's not like this curve skyrockets.” 32:37
“Even for a fixed model, we can get very different outcomes for these things, because they really are sophisticated engineered systems at this point.” 45:12
“The true cost of a token: the estimates vary wildly. For every dollar we spend, it could be as low as two and as high as 20.” 42:02
“The architecture everyone has arrived at, these stacked transformers that we make very deep and very large, are tremendously inefficient.” 12:00
“Experts display an augmentative style. They iterate with the AI. They push back. They complain. They change their requirements.” 47:24
Robohub · Ellen H. Rumley · 2026-09-09
The article sets out US defence funding, procurement rules, export controls and pending legislation that shape the market for robotics hardware.
“The Pentagon’s fiscal year 2026 budget requested an unprecedented $13.4 billion for developing autonomous systems, a major increase from years prior.”
“In the first quarter of 2026 alone, defense technology venture capital investments reached a record $19.8 billion.”
“Hardware inherently requires higher, long-term capital expenditures and yields lower short-term financial returns compared to software.”
“The U.S. mostly relies on foreign imports for baseline industrial robot installations.”
“Contrast this with Japan, which relies on foreign imports for a mere 2% of its domestic installations.”
“The challenge facing the U.S. does not seem to be a lack of federal funding, but rather how the funding is distributed across multiple agencies with different mandates.”
Ahead of AI · Sebastian Raschka · 2026-09-09
The reported reuse of transformer blocks in OpenAI's Astra raised a monitorability concern, and Raschka's argument is that looping neither shortens nor obscures the reasoning trace by itself.
“A Looped Transformer is essentially an architectural tweak, with the main idea being to pass the intermediate representations through the same transformer blocks multiple times (instead of just once).”
“The whole idea here is that we increase the effective depth from 22 to 44 block applications without adding another set of transformer weights.”
“Consequently, in KV caching, the resulting keys and values are also different between these two transformer stacks, just like in the no-looping case. So there are no KV cache-related savings either.”
“In my opinion, the success behind Astra, meaning its good modeling performance, is likely primarily due to other reasons, namely improved training recipes and training data.”
“Using fewer tokens could just mean that the model is more capable and makes fewer mistakes, uses less backtracking, and so on.”
“The researchers estimate that SMELT requires about 6.8-18% less training compute to reach the same validation loss within the studied compute range.”
What labs shipped 10
Robohub · 2026-09-25
A ROS-compatible, hardware-agnostic stack for control, motion planning and grasping lets developers change arms, grippers and sensors without rewriting drivers.
“Intrinsic (an AI robotics group at Google) announced that the company is making parts of its platform open source.”
“A hardware-agnostic, real-time control framework that delivers sensor-based control.”
“The idea is that users can swap out different robot arms, grippers, sensors, and more without needing to rewrite drivers.”
“This is a ready-made reference design for a real-world use case.”
The Robot Report · 2026-09-24
The largest SCARA maker enters lightweight collaborative arms designed for mobile manipulation and cleanroom use.
“The new robot arm has a 900 (35.4 in.) reach, a 6 kg (13.2 lb.) maximum payload, and an RC-A1010 controller.”
“The AX6 complies with ISO Class 5 (ISO 14644-1) cleanroom standards for controlled environments.”
“Yes, the lightweight and power design was specifically built with AMR integration in mind.”
The Robot Report · 2026-09-24
Closed doors kept inspection robots out of the parts of a plant that matter most, and the add-on extends ANYmal coverage without changing site infrastructure.
“When ANYmal approaches an unlocked door on its planned route, a radar sensor detects its presence and triggers a retrofitted dormakaba ED 100 or ED 250 mounted on the door.”
“Over several weeks of continuous testing, ANYmal has operated reliably, repeatedly requesting access, passing through controlled doors, and completing its inspection routes without manual intervention.”
“Because the robot logs every credential scan and transit with a precise timestamp, it creates a clear audit trail that provides better visibility than a manually propped-open door.”
“Fire doors compartmentalize almost every process facility to isolate hazards, placing 30% to 50% of critical equipment behind restricted barriers.”
Latent Space · 2026-09-24
An open model that predicts video and actions jointly, with LeRobot and Jetson support, lowers the entry cost for robot policy development.
“BFL released an open-weights 7B world-action model that takes #1 on RoboLab.”
“It beats the previous best open model by 6.1 points with 56% fewer parameters, and runs up to 3.95x faster.”
“It predicts video and actions jointly.”
“It ships with LeRobot integration and Jetson deployment; backbone and embodiment finetunes are open.”
The Humanoid Hub · 2026-09-24
A new German humanoid maker targets logistics sorting with in-house development, but the specifications come only from the manufacturer via two secondary sources.
“The P0 joins a growing European humanoid ecosystem.”
“Degrees of freedom: 32 in total, 10 of them in the hands.”
“They rely on their own technology development instead of licensed platforms from the USA or China.”
The Robot Report · 2026-09-23
Qualcomm is placing its robotics chips in the hobbyist-to-production pipeline through Arduino, which lowers the entry barrier for hardware founders prototyping robots.
“One of the blockers of the robotics industry getting into the AI era was that models, until just a couple of years back, were stuck on the cloud.”
“[Edge computing] is the only way to really fulfill many of the constraints of real-world deployments of these robots that might end up in warehouses or manufacturing sites that are offshore, where you cannot expect even close to a reliable high-speed connection to the internet.”
“It pairs a Dragonwing IQ-8275 processor, delivering up to 40 dense TOPS of AI performance, with 16 GB LPDDR5 memory, 64 GB eMMC, and expandable M.2 NVMe storage alongside a dedicated STM32H5 real-time microcontroller.”
Latent Space · 2026-09-22
A phone maker now leads open-weights models, and Xiaomi says it will open-source its RL environments and training recipes.
“MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46).”
“Large batches on a fully asynchronous architecture, with 1,568 samples per update, training at up to 1M context length, and 3.5 to 3.7B tokens per step.”
“Relative comparison within each group gives long-horizon RL tasks more precise and more diverse reward signals, closes a self-improvement loop, and steers the model toward shorter paths and fewer tokens per task.”
“@zephyr_z9 cites 130 hours, 75B tokens, and $2.6M for the RL run behind the result.”
“High-quality open RL environments may now be as strategically important as pretraining corpora were in the last cycle.”
The Humanoid Hub · 2026-09-21
Dexterous hands are a named bottleneck for humanoids, and a fully actuated 22-DoF hand at this price gives developers a component to buy separately from a whole robot.
“Compared with the predecessor Dex5-1, Unitree has added an extra roll axis on the little finger.”
“All 22 joints are motorized and backdrivable, each equipped with dual-encoder control and a torque-threshold protection function that prevents overload damage.”
“According to Unitree, the payload is around 2 kg at peak and about 1 kg in continuous operation per hand.”
“Unitree gives a starting price of around 39,900 RMB (roughly 6,500 US dollars) per hand in China, plus taxes and shipping.”
“It fits Unitree's own platforms (G1, H1) but is also designed for other humanoid systems and teleoperation setups.”
The Robot Report · 2026-09-19
The linear rail extends a cobot with a 175 cm reach along parts of 3.6 m to 6 m, and the automated seventh axis is planned for the end of 2026.
“A Hirebotics force- and power-limited robot has a 68.9-in. (175 cm) reach, but moving a 12-ft. (3.6 m) frame, a 20-ft. (6 m) structural assembly or a large weldment, for example, extends well beyond that envelope.”
“The linear rail turns that customer ingenuity into a standard Hirebotics offering.”
“Today, Beacon pauses the process when the cobot needs to move and prompts the operator to reposition it.”
Latent Space · 2026-09-19
An open dataset of this size, collected on an $8K bimanual setup, gives robot-learning teams training data and evaluation tooling without building their own hardware.
“ABC-130K is the largest open teleop dataset to date: 3,500 hours, 130K+ episodes, 195 tasks, collected on an $8K bimanual setup, with open hardware, training code, sim, and eval.”
“The full ABC release includes code, 400+ hours of sim data on 24 tasks, and 5,850 labeled policy-evaluation episodes.”
“The baseline science included a sim-to-real correlation of r = 0.91 on task progress and studies of offline metrics, scaling laws, and conditioning.”
The long listen 4
Machine Learning Street Talk · Frank Hutter · 2026-09-23
Frank Hutter argues that deep learning now beats CatBoost and XGBoost on tabular data because of TabPFN, a transformer that is pre-trained on hundreds of millions of synthetic datasets and executes a learning algorithm in one forward pass. The whole training table is given as context, and the network attends to it to predict the test rows. Hutter describes TabPFN as the first tabular algorithm learned end to end, with the cross-entropy loss on the held-out part of each meta-training dataset as its objective. He says it approximates the Bayesian posterior predictive distribution directly, skipping the posterior over parameters that MCMC or variational inference would compute. The prior is a distribution over structural causal models. Real tables, he says, are too few and too unlike each other to train on, so the synthetic data was necessary.
The host asks what causality means for such a model and raises the objection that coding agents such as Claude Code could build adaptive pipelines around XGBoost. Hutter agrees agents are strong for feature engineering, exploratory analysis and lookups such as national holidays, and says an agent should then call a tabular foundation model, which he says LLMs cannot match on numeric statistics. The host asks whether the model could be transductive or adapt at test time. Hutter names early exit, fine-tuning, prompt tuning and choosing between priors in a forward pass, but keeps the interface inductive: a test row must get the same prediction whatever batch it is in. On Google's TabFM he says it is larger and slow, and that it may be premature to go that big. He says that TabPFN did not focus on TabArena and that this deserves more attention. The episode also covers the AutoML history from SAT solvers to neural architecture search, relational data, the SAP acquisition and hiring, and a closing update on TabPFN 3.5.
TabPFN 1 appeared a week before ChatGPT in November 2022 and handled about 1,000 rows. Hutter gives 10,000 rows for version 2, 100,000 for 2.5 and one million for version 3, roughly two orders of magnitude a year, with 10 million as the next target. Version 2 alternated attention over rows and columns at cost n squared times m plus n times m squared. Version 3 adopted the TabICL architecture, quadratic in rows and n times m squared for embedding. A KV cache now separates training from prediction. Regression is treated as classification over about 10,000 bins. Hutter says a mean-and-variance head did worse in his student's experiments, and that the binned output can express two-peaked predictions. TabArena ranks methods by Elo scores, and a benchmark run costs about $2 on a GPU. Hutter states that TabFM is about 15 times slower per forward pass and fails on out-of-memory errors above roughly 100,000 rows. He says meta-training with interventions lets the model predict intervention effects from observational data, within limits of identifiability, and that some papers report smaller randomised trials with similar performance. The current models do not yet recover causal structure. A former student condensed a 1,000-row medical dataset into two rows with equal predictive performance. In the added update, TabPFN 3.5 took first place on seven benchmarks and was over 100 Elo points higher than TabFM while faster, or 20 times faster at equal quality. Nick Ericson reports first rank on a 2015 Kaggle competition in one line of code and one minute on one GPU, where the winner was a stack of 36 models.
“Until the day before yesterday, figuratively speaking, deep learning did not work for tabular data. Now it works dramatically better than CatBoost and XGBoost.” 10:08
Q “To what extent can we expect these models to be causal, and what would that even mean?” 1:18:49
“What we approximate is directly the Bayesian posterior predictive distribution.” 32:26
“LLMs do not work out of the box for this type of data. They are not made for tabular data.” 44:15
Q “So how does it compare to some kind of adaptive architecture like that?” 49:04
“You can actually meta-learn to make predictions about interventional effects from purely observational data, within some limits.” 1:23:32
Machine Learning Street Talk · Alexander Mattick · 2026-09-21
Alexander Mattick argues that a model trained with no prior information has to relearn everything from scratch, and that much of the current vocabulary in generative and world modelling names techniques that already existed. He calls energy based models not useful anymore as a way of thinking about most problems, because anything that predicts an unnormalised intensity map can be framed as one. For generation he names diffusion and flow matching as the better methods. He says JEPA and world models are ultimately branding: JEPA-like models existed before JEPA, and Schmidhuber's world model paper was model-based reinforcement learning. On Rich Sutton's reward is enough, he says the statement is true in the limit, since a reward can be written for any task, but that finding such a reward and optimising it may not be realistic. He adds that he does not want to put words in Sutton's mouth. His alternative is constrained reinforcement learning, which encodes known constraints and prior data instead of tuning reward weights.
The host presses first on energy based models, citing Yann LeCun's view that being non-generative is a good property. Mattick disagrees with the premise: an energy model placed in an MCMC sampler produces samples, and the question is whether that is economical, which today it is not. He concedes two uses, pairwise ratios and physics settings where an energy is known. Asked about LeCun's companies AMI and Logical Intelligence, he doubts that what they build in five years will be energy based models as he understands the term. He also concedes that inference is an overloaded word he must use carefully. On theories of deep learning, the host lists geometric, computability, complexity and statistical physics views. Mattick prefers whichever makes falsifiable predictions and says no current approach can tell an implementation bug from a failed theory. On the manifold hypothesis, after the host describes Goodfire's geometric structures, he calls it neither true nor false nor a useful framing and says interpretability results rely on simplifying assumptions. Asked whether networks can have hard constraints, he doubts it in general and concedes that a fair comparison of constrained and unconstrained learning is hard to design. The host also raises a robot demo of physical prompting, and Mattick says a nonlinear model predictive controller could behave the same way and that the open question is guarantees.
Mattick traces inference from variational inference, which fits a single normal to bimodal data, through rejection sampling, which scales badly with dimension, and MCMC, which is slow. Energy based models train without sampling but need MCMC to generate. Diffusion adds a time variable and follows a stochastic differential equation. Flow matching fixes a reference path and trains by regression, without solving an ODE at every step. On theory he covers neural tangent kernels, mean field theory and Barron's dimension-free convergence bound. On constrained reinforcement learning he lists CPO, projection methods, Lagrangian methods that oscillate, and local optimisation methods. He describes OpenAI Five's reward table, in which killing an enemy hero has a negative weight, and says an explicit constraint avoids retuning correlated weights. He cites a paper showing that uniform accuracy for ReLU networks needs exponential data. He separates model-based reinforcement learning, where data collection is in the loop, from a world model, where it is not. He says a perfect world model of chess would not make someone Magnus Carlsen. He says a robot that succeeds 90 percent of the time would be refused at BMW. He also reports losing two or three months to an implementation bug.
“I think that energy based models are not really useful, or at least not useful anymore.” 39:49
Q “Are there any situations, do you think, where EBMs are better?” 53:08
“The question is more whether this is a realistic assumption to make: whether having a god-given reward function is realistic, even finding it, and once we have it, whether it is realistically optimizable with any technique.” 1:30:33
“We still have the problem of now needing to technically relearn everything from scratch.” 1:39:09
“There were JEPA-like models before JEPA, and there were world models before world models.” 2:03:31
“I think they are set up to look very nice in demos and to do some really impressive stuff. But how do you know how many times it failed to do something?” 2:13:27
Machine Learning Street Talk · Tom McGrath · 2026-09-02
Tom McGrath of Goodfire argues that interpretability is a natural science done entirely on a computer, and that it should be used to steer training as well as to read models. He describes current training as close to open-loop control. In reinforcement learning with verifiable rewards (RLVR) the model receives a binary success or failure signal and goes wherever the data takes it. Interpretability gives a readout of what a forward pass computes and of the directions in which a backward pass will change the model, which he says allows closed-loop control. He calls this intentional design and names controlled generalization as its first step: taking some things from a dataset and not others. His second argument is geometric. Many concepts, such as days of the week and months of the year, are represented as curved manifolds, and a sparse autoencoder (SAE), whose inductive bias is that features lie on rays from the origin, breaks them into pieces. He offers stepping off the manifold as the explanation for why activation steering sometimes works and sometimes produces gibberish.
The host presses first on the safety community's "forbidden method", the use of interpretability signals in training. McGrath accepts the underlying concern, that training against a monitor also trains away the ability to monitor, but says a small vocal minority has turned it into a taboo and that most active practitioners regard the research as reasonable. He agrees that back-propagating through a probe is a bad idea. The host then asks whether shaping training could make models degenerate or less adaptable, and whether concept ablation can win against stochastic gradient descent. McGrath answers that it can, but that it will be hard and needs a new science and a new engineering discipline. The host also raises the tension with the bitter lesson. McGrath replies that nothing in the process is specified by humans, and that he disagrees with Richard Sutton about scalar rewards: reward might be enough in principle, but is not enough today. On multi-agent systems, collusion and red teaming he agrees with the host's worries and says he has no particularly good solution. On Neel Nanda's blog post he says he disagrees, has longer timelines, and thinks SAEs remain useful although the manifold approach fits what networks do better.
For readout, McGrath uses an SAE as a gradient-attribution tool: the gradient at the SAE layer is dotted with the decoder and multiplied by the activations. On pirate-speak maths data this shows pirate features, and he calls it a crude approximation of the true parameter update. Positive preventative steering clamps a persona direction upward during the forward pass, which he compares to holding a radiator next to a thermostat, so that learning in that direction is neutralised. Inoculation prompting does the same in text. In the hallucination work, a probe amortises a model with web search, and in-context interventions alone reduced downstream hallucinations. He gives three hypotheses for why models hallucinate while knowing better: checking occurs earlier in the network than generation, reinforcement has been insufficient, and inventing facts is useful in fiction. Predictive data debugging clusters DPO pairs by SAE feature deltas. For geometry, an Ising model fitted to feature co-activations recovers splines, and block-sparse featurizers learn subspace sizes adaptively. A modular-addition study found Llama 3.1 8B using a base-10 Fourier operation through a shared addition module, with similar evidence in Llama 70B and DeepSeek V4 Flash. Unpublished work on Gemma 31B with a weak grader found comments that deceive the grader, and vectors built from synthetic data fire on them.
“Interpretability is the thing that lets us go to closed-loop control, because we can say we are going to go in this direction.” 15:03
“If you back-propagate through the probe you are just cooked. This is basically always a bad idea.” 24:30
Q “Do you think in principle we can fight against SGD and make this successful?” 27:53
“What we show in this paper is that a lot of these representations route through a general addition module.” 1:14:17
“It is a bit of an indictment on the field that we cannot yet do it.” 1:24:48
Q “How do you think that self-conceptualization actually emerges?” 1:30:29
Welch Labs · 2026-09-01
The narrator of this Welch Labs video argues that a single change, adding each layer group's input to its output (a skip connection), removed the depth limit that stalled deep learning in early 2015. The change is the residual network, or ResNet, from Jian Sun's group at Microsoft Research Asia, described in a 12-page paper of December 2015 that the narrator says became the most cited paper of the 21st century. The narrator gives the diagnosis as gradients that become unreliable as depth grows, the "shattered gradient problem" named in a 2017 paper. The narrator then argues that the skip connections created a continuous path from input to output, later called the residual stream, that every layer adds to, and that this stream works as a working memory for the model.
The video is a single-voice explainer with no host or guest, so nothing is disputed and no position is conceded. It first shows the puzzle: deeper models that should be able to copy a shallower one and yet score worse. It then rebuilds the effect with an eight-layer convolutional network and loss landscapes, introduces the fix, and follows the reinterpretation through a 2016 Cornell study, the 2017 transformer and a 2023 Meta study of vision transformers. The narrator compares this sequence to Planck's quanta leading to quantum mechanics, and says it remains to be seen whether most of the discoveries of this wave of AI have already happened. A sponsor segment on Jane Street describes Alok Perinic's blog post on positional encodings, which finds the space of possible encodings constrained, the most sensible ones already in use, and some approaches unexplored. The video closes with a Patreon appeal.
Jian Sun's team trained 30 layers with He initialization, after a Xavier initialization left the error at 100%. The 30-layer model reached a 16.59% error rate, against 13.34% for a 14-layer model. The team noted that adding 16 identity layers to the 14-layer model would give the deeper model the same output. In the narrator's own runs on ImageNet, accuracy was 44.1% at eight layers, 56.7% at 14, 62.5% at 20, 63.6% at 26, 62.8% at 34, 56.6% at 56 and 38.9% at 74 layers. Adding skip connections to the 74-layer network raised it to 72.6% and smoothed the loss landscape of its early layers. Kaiming He is shown presenting networks of over 150 layers. The Cornell team found that removing one layer from a 56-layer ResNet barely changed performance, while the same removal damages non-residual networks such as AlexNet. A separate team later showed evidence that residual networks do both: learn hierarchical representations in subsets of layers and refine the residual stream. The Meta team found large activations at a few unimportant positions in DINOv2, which has 40 layers and a 37 by 37 by 1536 residual stream. Classifiers trained on those positions scored 85.2% on the cars data set, against 10.8% for ordinary positions. Adding learned register tokens, which are discarded at the end, removed the large activations.
Q “If the 30-layer model was capable of achieving at least the same performance as the 14-layer model, why couldn't the team's optimizers find these solutions?” 2:59
“Between each pair of layers in the model, simply take the input activation tensor and add it to the output of the layers.” 17:02
“This flow of information from the input to the output of our residual network, iteratively refined by each layer, was later given the name the residual stream and has become one of the defining features of modern AI.” 26:08
“The model learns to recognize patches containing little useful information and recycles the corresponding tokens to aggregate global image information while discarding spatial information.” 29:21
“These results strongly support the view of the residual stream as a working memory for the model.” 31:46