用3D世界感知提升机器人泛化能力,少用真实数据也能高效训练。
WSA$_1$: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control

- 基于3D世界-空间-动作联合建模,让机器人理解行为对环境的因果影响。
- 仅用6000小时演示数据(含1000小时真实数据)即达93%操作成功率。
- 适合追求低成本、高泛化能力的机器人系统研发团队使用。
近年来,具身AI的发展使机器人基础模型(RFM)成为通用机器人系统主流方法。通过在大量机器人示范数据上进行模仿学习,RFM已能将视觉观察和语言指令映射为连续机械臂动作。然而,现有RFM缺乏对物理动态及行为因果效应的内在推理能力,导致2D视觉感知与3D具身交互之间存在根本性错配,严重限制其在真实任务中的泛化性能。为此,我们提出WSA₁,一种基于新型3D-Centric World-Spatial-Action建模范式的全新RFM。该模型不仅能学习未来行为相关的3D世界感知表征,还能建模3D世界状态变化与机器人动作之间的相互约束,从而增强行为泛化能力。值得注意的是,WSA₁仅需6000小时专家示范数据(其中仅1000小时来自真实机器人)即可完成高效预训练,在RoboTwin2.0仿真基准上实现93%的成功率,并在真实机器人任务中相较当前最优RFM平均提升20%性能。结果表明,结合3D-centric世界-动作联合建模,无需大规模真实数据即可实现可泛化的RFM,为通用机器人系统提供了一条可行且经济的路径。
原文摘要 · Abstract (English)
Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By leveraging imitation learning on extensive robot demonstrations, RFMs have achieved impressive capabilities in mapping visual observations and language instructions to continuous robotic actions. However, current RFMs lack an inherent ability to reason about physical dynamics and the causal effects of robot behaviors on the 3D physical world. This creates a fundamental mismatch between 2D-centric visual perception and 3D-centric embodied interaction, severely limiting the generalization ability of RFMs in real-world tasks.To address this gap, we present WSA$_1$, a novel RFM built upon proposed 3D-Centric World-Spatial-Action modeling paradigm. It not only learns 3D world-aware visual thought for future robot behaviors, but also models mutual constraints between 3D world state transitions and robotic actions to enhance behavior generalization. Notably, WSA$_1$ achieves highly data-efficient pre-training with 6k hours of expert demonstration data (only 1k hours from real robot), while delivering competitive manipulation performance (93% success rate) on RoboTwin2.0 simulation benchmark and achieving +20% average boosted performance over state-of-the-art RFMs on real-world robot control tasks. These results reveal that generalizable RFM can be attained without large-scale real robot data when paired with 3D-centric world-action joint modeling, which offers a practical and affordable pathway to generalist robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。