Reward AI Avoids Robot Data, OM-1 Zero-Shot "Body Swapping": Smooth Packaging and Cable Unplugging with Two Hands

09/16 2026 404

Smart, smooth, and dexterous—a U.S. startup offers a new solution for cross-embodiment generalization in robots.

Just 24 hours ago, U.S. robotics startup Reward AI unveiled its first foundational robot model, OM-1. It features three key characteristics: learning directly from human operation data without using teleoperation or robot-specific data; zero-shot transferability to tabletop robotic arms, industrial robotic arms, and humanoid robots using the same model; and continuous execution of tasks like packaging, bartending, and folding clothes at speeds close to human performance.

Most critically, there's no need to restart training from scratch when switching to a different robot. This is OM-1's most impressive capability.

Reward AI also released a set of demonstrations: a robotic hand pressing down on an RJ45 clip to unplug a locked Ethernet cable; two robotic arms collaborating to package a smartphone; a dual-arm robot folding clothes, retrieving a liquor bottle, and mixing drinks; OM-1 can drive a humanoid robot for full-body operations and navigation, though further details remain undisclosed.

From the demos, all clips are shown at original speed, with long-duration tasks like smartphone packaging, bartending, and clothes folding completed in under 30 seconds. NVIDIA Robotics lead Jim Fan's brief comment on the release: "So smooth!"

But what OM-1 truly aims to transform is the relationship between robot models, training data, and hardware.

01 Cross-Embodiment Isn't New—Eliminating Robot Data Is Key

Controlling different robots with the same model isn't a goal first proposed by OM-1.

Physical Intelligence's 2024 foundational robot model π0 used open robotics datasets and data collected from eight robot types. Google DeepMind's Gemini Robotics 1.5 demonstrated motion transfer between ALOHA 2, dual-arm Franka, and Apollo humanoid robots; the July 2026 release of Gemini Robotics 2 claimed adaptation to new dual-arm robots using just hours of data, typically under 200 samples.

OM-1 takes a more radical data approach: the same model responsible for operational decisions learns solely from human demonstrations, eliminating the need for robot-specific training data collection or fine-tuning.

This path stems from research by Reward AI's two founders. CEO Zipeng Fu holds a PhD in Computer Science from Stanford, mentored by robotics learning scholar Chelsea Finn, with contributions to projects like Mobile ALOHA and HumanPlus. CTO Chen Wang, also a Stanford Computer Science graduate under Li Fei-Fei and C. Karen Liu, focuses on human motion learning, dexterous manipulation, and long-term planning.

OM-1 directly extends Chen Wang's DexCap research at Stanford: instead of remote-controlling robots first, it records how humans naturally complete tasks, then transfers this experience across different robots using the same model.

When collecting data via teleoperation, human actions pass through a specific robot's joints, grippers, and control systems, inadvertently incorporating hardware-specific traits. OM-1 shifts data collection upstream: humans operate naturally, the model learns reusable skills, and different robots' low-level control systems handle execution.

02 Capturing "Subconscious Bodily Intelligence"

When people pick up a cup, unscrew a cap, or fold clothes, they rarely calculate each finger's movement. Decelerating before contact, increasing grip when an object slips, or nudging a misaligned box lid—these pushes, slides, twists, and force adjustments constitute what Reward AI calls "subconscious bodily intelligence."

To capture this, the team developed the wearable Omnibody Hand. Rather than mechanically replicating all human hand joints, it uses seven degrees of freedom to preserve key functions: precise pinching with thumb and forefinger, coordinated multi-finger force application, and switching between precise and power grips.

The device simultaneously records images, hand trajectories, inter-finger distances, tactile signals, and force magnitudes. The model learns not just "where the hand moved" but also when contact occurred, whether an object was securely grasped, and the force needed for pulling or lifting.

To track rapid movements, Reward AI added electromagnetic tracking alongside visual-inertial methods. In tests across eight speeds (ten trials each), their approach reduced average overshoot error at maximum speed from 24.9mm to 9.5mm—a ~60% improvement. While this measures hand-tracking accuracy, not robot task success, it shows the system can record demonstrations accurately without requiring humans to slow down deliberately.

For a previously unlearned task, Reward AI claims OM-1 requires under 30 minutes of total human demonstration data. New tasks still need demonstrations; "zero-shot" primarily means the same model doesn't require robot-specific fine-tuning when transferred to another robot body.

03 From Unplugging Cables to Quad-Arm Packaging: The Model Must Adapt in Real Time

OM-1's demonstrated tasks go beyond simple grasping and placing.

Unplugging an Ethernet cable exemplifies this. The robotic hand must precisely locate the RJ45 clip, press it down, and pull outward. A slight positional error prevents release; excessive force risks damaging the connector.

Opening a refrigerator door tests force adaptation, as the robot must adjust to unknown resistance. Bartending and clothes folding involve multiple sequential steps, while smartphone packaging requires four robotic arms to coordinate around a single object.

Reward AI notes several emergent behaviors in demonstrations: when one robotic arm deviates, others adjust their movements to compensate; the model retries after failures; it alters actions when objects are externally disturbed; and if environmental changes exceed thresholds, the robot halts execution.

These behaviors indicate OM-1 doesn't generate fixed trajectories played end-to-end. Instead, it continuously adjusts based on contact states, task progress, and other arms' actions.

OM-1's true radicalism isn't being the first to propose cross-embodiment but attempting to remove robot hardware from upper-layer model training data: making human demonstrations a long-term reusable data source, with specific robots merely serving as different bodies executing these skills.

If this approach withstands prolonged operation and diverse hardware validation, robot data may no longer depreciate with each hardware generation. Human operations recorded today could teach robots manufactured tomorrow.

That's the most noteworthy gamble behind OM-1's release.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.