arXiv:2608.27033cs.RO2026-08

一个能同时执行动作和模拟世界演化的机器人模型,让机器人的经验可迁移。

Riemann-1.0: An Embodied World Action Model for Physical AI

论文配图:Riemann-1.0: An Embodied World Action Model for Physical AI
图 1 · 摘自论文原文
  • 用统一因果序列建模视觉、状态和动作,实现动作与世界演化的联动。
  • 在真实世界任务中达85%成功率,长序列任务领先开源基线15%。
  • 融合人类视频、示范数据和机器人轨迹,实现跨模态经验迁移。

我们提出Riemann-1.0,一个全因果自回归的世界动作模型(World Action Model),用于具身智能。该模型统一建模多视角视觉观测、机器人状态与具身特有动作,以因果状态转移形式表示机器人动作与世界演化。不同于现有基于联合生成、视频先行预测或解耦建模的WAM,Riemann-1.0将在线机器人策略执行与动作条件下的世界仿真统一于单一模型,既可作为可执行策略,也可作为多具身视觉世界模拟器。为在异构数据源间扩展具身经验,我们进一步构建渐进式具身预训练框架,统一人类第一人称视频、手持机械臂示范及异构机器人轨迹,在共享世界动作建模目标下学习。基于20万小时以上交互数据,Riemann-1.0逐步将大规模具身经验转化为可执行的机器人操作能力。在仿真与真实世界任务中均达到领先性能:在RoboTwin2.0上成功率达94.3%,在LIBERO上达99.0%,在长程组合任务RoboCasa-365上达62.6%,较前最优方法提升8.4%;在长程真实世界操作任务中,成功率为85.0%,进展成功率94.4%,优于最强开源基线15%。结果表明,统一世界动作建模与渐进式具身预训练能有效将大规模具身经验转化为通用机器人操作能力。

原文摘要 · Abstract (English)

We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unified causal autoregressive sequence, representing robot actions and world evolution as causal state transitions. Unlike existing WAMs based on joint generation, video-first prediction, or decoupled modeling paradigms, Riemann-1.0 unifies online robot policy execution and action-conditioned world simulation within a single model, enabling it to function as both an executable robot policy and a multi-embodiment visual world simulator. To scale embodied experience across heterogeneous data sources, we further develop a progressive embodied pretraining framework that unifies learning from egocentric human videos, handheld-gripper demonstrations, and heterogeneous robot trajectories under a shared World Action Modeling objective. Built upon 200K+ hours of interaction data, Riemann-1.0 progressively transfers large-scale embodied experience into executable robot manipulation capabilities. Riemann-1.0 achieves state-of-the-art performance across both simulation benchmarks and real-world manipulation tasks. It achieves success rates of 94.3% on RoboTwin2.0, 99.0% on LIBERO, and 62.6% on the long-horizon compositional benchmark RoboCasa-365, outperforming the previous best method by 8.4% On long-horizon real-world manipulation tasks, Riemann-1.0 achieves a Success Rate (SR) of 85.0% and a Progress Success Rate (PSR) of 94.4%, exceeding the strongest open-source baseline by 15% in SR. These results demonstrate that unified World Action Modeling together with progressive embodied pretraining effectively transforms large-scale embodied experience into generalizable robot manipulation capabilities.

具身智能世界模型机器人操作因果建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。