09/30 2026
486

Author|Mao Xinru
In early September, OpenAI released GPT-6 Astra.
The next day, independent evaluation agency Robocurve connected it to a dual-armed robot with a simple task: pick up building blocks and place them in a bowl.
It succeeded 19 out of 20 times, consuming only about one-sixth of the tokens used by the comparative model, Fable 5.1.
Immediately, discussions in the embodied AI community exploded. Claims that general-purpose large models would take over robotics and that embodied AI would become a mere subset of foundational models spread rapidly.
But the cooldown came quickly. The RoboDojo team subsequently concluded in their official real-world evaluation: Astra repeatedly issued physically unreasonable or unsafe actions, with some incidents even causing hardware damage, forcing the evaluation of 18 planned tasks to terminate early.
Three weeks later, the RoboDojo leaderboard saw two new leaders: Simate, a company founded just three months earlier, and "MMLab HKU & Chaowei Power KAI."
Both ascended to the top with Physical RSI.
RSI (Recursive Self-Improvement) is rapidly gaining popularity (popularity) from the large model domain into the robotics industry.
Anthropic revealed in a June report that over 80% of the code merged into its codebase was written by Claude, labeling RSI as the ultimate form of self-evolution.
In August, Wang Xingxing unveiled a physical AI self-evolution system at the World Robot Conference; in September, Chen Long, former head of Xiaomi's XLA large model, ventured to Silicon Valley to pursue the same direction: a self-evolving embodied brain.
Suddenly, large model companies, robotics firms, and top talent all seemed to be talking about RSI.
But is this just a buzzword marketing campaign or a vanguard battle for a paradigm shift in embodied AI?
Demystifying Embodied RSI: It's a Goal, Not a Technology
RSI is not a new concept.
Its origins trace back to AlphaZero's self-play, where AI surpassed top human chess players through internal game theory without relying on human game records.
In 2025, DeepMind clarified the logic: human data was nearly exhausted, and AI would now learn from its own experiences.
Subsequently, Karpathy (one of OpenAI's founding members) demonstrated AutoResearch, showcasing a minimal closed loop where AI conducts research to improve AI, followed by DeepMind's Dream-RSI paper and Zhipu's minimal closed-loop validation.
These efforts spanned games, large models, and infrastructure optimization but shared a common direction: enabling systems to learn from their own experiences rather than relying on human-provided data.
However, the hotter a concept becomes, the more prone it is to misuse.
Fundamentally, RSI is a goal, not a technology.
Reinforcement learning, self-supervised fine-tuning, and continual learning are specific technical paths with precisely definable elements: rewards, ground truth, and gradient updates.
RSI, by contrast, refers to the system's ability to continually improve using its own experience—the architectural implementation remains open.
An industry-standard validation method involves examining the three words in Recursive Self-Improvement individually. Improvement is the easiest to achieve: data reflow (feedback), training, and performance gains suffice.
Next is Self: if data collection, task selection, and parameter adjustments remain highly human-dependent, the system merely represents human-assisted automatic optimization. The hardest is Recursive, as it requires not only stronger capabilities but also the mechanism for self-improvement to be improvable itself.
The large model domain is considered to have touched RSI because AI R&D is now accelerating AI R&D. Claude writes Claude's code; AutoResearch optimizes AutoResearch's processes—a true recursive closed loop.
By this strictest standard, the ultimate embodied RSI would involve a robotic system improving itself: upgrading hardware when needed or even assembling a new robot.
No one has reached this stage yet.
Meanwhile, academia attempts to measure RSI.
A recent joint review by Shanghai Jiao Tong University, Tsinghua University, ByteDance, and the Shanghai AI Laboratory, titled *The Last AI Built by Humans*, systematically categorizes RSI into five levels for the first time:
L1: Execution autonomy—humans design tasks and strategies; agents execute and update.
L2: Strategy autonomy—agents autonomously design and update strategies.
L3: Experience acquisition autonomy—agents autonomously plan experience acquisition.
L4: Environmental adaptation autonomy—humans set boundaries; agents autonomously adapt.
L5: Recursive inheritance autonomy—agents improve the improvement system itself.
Although framed for general AI, this framework offers reference value for observing embodied RSI.
When evaluating existing embodied AI RSI projects against this scale, all remain in the weak RSI zone of human-machine co-driving, between L2 and L3.
AI can select strategies and proactively pull data for validation, but determining why a strategy is superior or which logic to rewrite for true Upgrade (upgrades) still requires human decision-making.
Between L2 and L3, teams have chosen vastly different paths.
Some pursue RSI at the execution layer, having agents solidify reasoning into code to build reusable skill trees. Others focus on RSI at the strategy layer, enabling robots to perform online reinforcement learning during real execution, updating gradients only for actually executed actions. A third group pushes RSI to the R&D layer, having AI agents handle model development, data processing, and experimental evaluation.
They share the same goal but answer "how to do RSI" in starkly different ways.
Three Engineering Paths, Three Flavors of "Self-Evolution"
Embodied RSI remains in its chaotic phase, with no unified industry standard.
We examine three distinct angles—task execution layer RSI, policy parameter layer RSI, and algorithm R&D layer RSI—to understand how different routes and tiers converge on the same goal.
The "tier" distinction here refers not to depth or R&D stage but to where self-evolution acts within the system.
If robot capabilities are broken into five layers—task understanding, task planning, skill generation, action generation, and action execution—these three embodied RSI cases target different rings, modifying different objects.
Stardust Intelligence's SmoothRL represents the execution-end attempt in this chain.
It addresses a specific engineering question: When a robot's executed action falls short of expectations, can it immediately update its Policy (strategy network) using real-world feedback instead of returning to the lab for offline retraining?
SmoothRL's core embeds Online RL into the asynchronous inference process of real robot execution.
In real-world deployment, robots output Action Chunks to their mechanical bodies. While executing these actions, the large model concurrently computes the next sequence in the background.
This creates a conflict: parts of the previous Action Chunk may be overwritten by new inferences before hardware execution completes.
Traditional Online RL incorporates the entire Action Chunk into loss calculations, training on many unexecuted actions and polluting learning signals. SmoothRL's key innovation distinguishes which actions within a chunk were physically executed versus overwritten, using only truly executed actions for Policy updates.
Instead of waiting for task completion to return data to the lab for offline retraining, SmoothRL creates a fast loop: execute, obtain real physical feedback, update the policy, and re-execute.
In Stardust's view, RL inherently forms a recursive cycle, and Online RL's strength lies in speed: the model updates strategies based on fresh real-world feedback.
This contrasts sharply with traditional offline training by drastically shortening the link between real feedback and policy updates.
If pursued long-term, real-world data could eventually reflow to foundational models, enabling longer-term capability accumulation. A Policy update for one task might serve not just that task but also as data and experience for the next foundational model iteration.
Thus, experience compounding occurs: after each online update, the model knows it has a stronger version, starting the next learning cycle from a new baseline.
Physical Intelligence's π0.7 follows a similar iterative logic, evolving continuously from π0.6 and historical reinforcement learning data.
SmoothRL fundamentally alters how robotic Policies update from the real world.
Qomolo Intelligence's RoboRSI focuses on task methods and skill accumulation.
It tackles the question: When robots struggle with unfamiliar tasks, can they autonomously develop better completion methods and precipitate (accumulate) them as reusable skills instead of requiring human engineers to handwrite control code each time?
RoboRSI doesn't entrust everything to a unified large model. Instead, it decomposes roles into distinct Agent types—Manager, Planner, Engineer, Reviewer—isolating task reception, planning, code execution, and retrospective modification into separate contexts. This allows different Agents to handle defined work while preventing individual Agent contexts from exploding in complexity.
This decomposition automates more processes and redistributes complexity between humans and Agents: Agents handle detailed task decomposition, code execution, and failure retrospectives, while humans receive structured results from Agents for easier decision-making.
The project's critical Knowhow centers on TSR (Top-down Skill Refinement), which addresses where Agents should modify and what to preserve during self-iteration.
A matching (supporting) skill trust mechanism manages the entire skill tree.
Skills verified as stable through real-world testing are marked as trustworthy, preventing Agents from repeatedly modifying validated modules in subsequent iterations.
Early in the project, the Qomolo team tried a bottom-up approach: having Agents first write low-level Skills before combining them into complete tasks. They quickly found Agents easily became stuck in local problems—essentially, " Split hairs " (obsessing over minor issues).
If one module failed, the Agent endlessly modified it, causing cascading impacts on other modules and further modifications. Though local code grew complex, it didn't necessarily advance the overall task.
They shifted to a top-down approach, first defining the complete task's topological structure and clarifying input/output interfaces and dependencies between modules before filling in specific Skill implementations.
Stable, real-world-verified Skills are protected via a Trust mechanism, preventing Agents from altering validated code modules in subsequent iterations.
A crucial cognitive shift occurred: the challenge of self-improvement isn't letting AI modify endlessly but knowing what to change and what to leave untouched.
This distinguishes RoboRSI from ordinary Agents.
Agents attempt task solutions; Reviewers handle high-weight reasoning for failure retrospectives; Skills accumulate mature capabilities; Contexts manage all historical experience.
The ultimate goal is to transform a lucky success into a mature capability directly usable for similar future tasks.
Simate has chosen to focus its research perspective on the more upstream dimension of algorithm development.
Its AutoResearch is no longer confined to motion adjustments or task skill generation during the robot execution phase. Instead, it directly embeds Agents into the entire process of robot algorithm development, attempting to use AI to undertake a significant amount of repetitive research work originally performed by algorithm researchers.
A distinction is made here: The Agent in RoboRSI serves as a robot task engineer, producing control code for the robot to perform tasks. In contrast, the Agent in AutoResearch acts more like an assistant to algorithm researchers, generating models, experimental configurations, and algorithm solutions.
Within the closed loop of AutoResearch, human researchers firmly control the research direction, objectives, and constraint boundaries, while the Agent undertakes a substantial amount of experimental execution work.
This ultimately forms a complete cycle: proposing research hypotheses, planning distributed experiments, modifying models or configurations, scheduling and executing experiments, invoking RoboDojo for evaluation, analyzing experimental results, and iteratively generating the next round of hypotheses.
To enable the Agent to genuinely modify and combine robot model components, Simate has made the SiPAI pluggable model framework publicly available. This framework aims to allow different model components, such as VLA, WAM, and World Model, to be understood, disassembled, replaced, and combined by the Agent.
This, in turn, allows research ideas to be directly translated into operational model experiments without the need to rebuild the entire engineering code system from scratch each time.
However, this system also falls within the weak RSI category of human-machine co-driving.
AutoResearch's evolutionary target is the robot's algorithms and model solutions, rather than directly enabling the physical robot itself to self-iterate in real-world environments.
It heavily relies on simulation evaluation infrastructure, allowing for a high concurrency of experiments in simulated environments. However, once deployed on real robots, environmental resets and irreparable physical failures remain significant obstacles.
Therefore, within today's discussion framework of embodied RSI, there is no so-called technological convergence.
Some are researching how to improve robot motions autonomously, others are exploring how to enhance methods for robots to complete tasks, and still others are investigating how to automate the research process for robots themselves.
Ultimately, they all point to the same question: Can the robot's next capability no longer be entirely reliant on human engineers to manually create it?
The physical world doesn't have a 'reset' button.
Regardless of which of the three paths is taken, they all encounter the same issue: Once in the physical world, many things that seem straightforward in the digital realm become much more challenging.
Failed experiments in software can be rolled back by reverting code and rerunning models; robots making mistakes in simulations can simply be reset to start over.
Real robots are not so convenient. If a robot knocks over a trash can, the environment changes; if it spills a cup, the next experiment faces a different scene.
The Qiomo Intelligence team even mentioned that in real robot environments, resetting can sometimes be more challenging than completing the task itself.
This is also one of the biggest differences between embodied RSI and software Agents: Every time a robot tries and fails, it may genuinely alter the world.
Therefore, the real challenge is not getting the Agent to try multiple times but enabling it to understand which attempts are worth keeping, which failures should be corrected, and whether the corrections have indeed improved the situation.
The Reviewer in RoboRSI and the Trust mechanism for validated Skills essentially address this issue: not allowing the system to modify infinitely but enabling it to discern what should be changed and what should remain untouched.
Unified simulation-real machine evaluation platforms like RoboDojo are also filling in the same piece of the puzzle.
Currently, the official platform has set up 42 simulation tasks and 18 real-machine tasks, with unified hardware, scene resets, and evaluation protocols to enable repeated verification of robot capabilities under relatively consistent conditions.
For RSI, this is particularly crucial. Without a reliable evaluation mechanism, there can be no reliable improvement.
If the system cannot determine whether it has improved, self-improvement can easily degenerate into continuous trial and error.
Therefore, what embodied RSI truly needs to establish is not just an Agent that can modify its own code but a complete chain: execution, feedback, evaluation, memory, update, and re-execution.
As this chain progresses, the role of humans will genuinely change.
Today, humans still need to decide goals, design experiments, and judge results; in the next phase, they may only need to set goals and boundaries; further down the line, the system may even need to decide for itself what should be changed in the next round.
This touches upon the core concept of RSI—Recursive.
Faced with the increasing number of RSI concepts today, to assess how far a system has progressed, one can consider three questions: Can the improved capabilities be retained and applied to other tasks? Has the iterative capability itself been improved, not just the task success rate? How much human involvement remains—are they still writing code and adjusting parameters, or have they stepped back to setting goals and safety boundaries? These three questions may be more important than the RSI label itself.
Systematic quantitative verification of Qiomo Intelligence's RoboRSI in simulated environments. Continuing along this path, robots in the future will need to answer not just how to perform a task but why they didn't perform it well, where they should improve next, and if a method consistently fails, whether they should switch to a different learning approach.
From SmoothRL's online updates to policies, to RoboRSI's reconstruction of Skill and Agent workflows, to Simate's attempt to integrate AI into the robot development process, these three cases are approaching the same question from different directions: Can a robot's experience become the starting point for its next improvement?
Clearly, this has not yet been achieved.
Resetting in the real world, evaluation, hardware differences, data costs, and the role of humans in goal setting and critical judgments have not been fully resolved.
Embodied RSI is still in its early stages, and more real-machine verification remains an unavoidable hurdle.
However, changes are already underway. In the past, the basic unit of robot development was a round of training, a deployment, or a new skill.
Now, more and more teams are asking: After a round of self-improvement, can the robot continue to improve itself from a new starting point?
If one day, robots not only are trained by humans but also begin to participate in deciding what they should learn next, how they should learn it, and how to verify whether they have learned it correctly, then embodied intelligence will truly begin to touch upon the meaning behind the term RSI. And this is not just a technological leap.
If robots can transform more and more of their training, debugging, and deployment work into their own growth process, the time and cost human engineers need to invest may decrease accordingly.
For robots that frequently change scenes and tasks, this means that the mode of relying on manual redevelopment and debugging each time may be replaced by a continuous self-iterative approach.
Self-evolution not only opens up the upper limits of robot capabilities but also presents another commercial possibility for the large-scale deployment of robots.
This may be what truly constitutes 'growth' for robots.