RoDyn让机器人模型更懂空间互动,提升抓取成功率42%。
RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- 用2.5D几何潜空间建模环境动态,融合视觉与空间先验。
- 通过掩码引导自回归架构,聚焦机器人与物体交互区域。
- 在真实世界模仿学习中成功率达42%提升,性能超越纯2D模型。
学习型世界模型作为机器人操作的神经仿真器具有巨大潜力。然而,主流的2D视频模型天然缺乏物理交互所需的时空与运动推理能力。本文提出RoDyn,一种新型的机器人-动态2.5D世界模型,在高效、几何感知的潜在空间中建模环境动态。通过提出的机器人-动态分词器,我们利用以RGB为主导的交叉注意力机制,显式耦合语义视觉外观与空间及代理中心先验,并引入动态掩码引导。此外,将这些掩码先验直接注入序列转换过程,使掩码引导的自回归架构专注于活跃的机器人-物体交互区域。大量实验表明,RoDyn在大规模数据集上实现了生成保真度的最新水平。关键的是,其预测能力显著提升了下游任务表现,加速了基于模型的强化学习,并在真实世界模仿学习中相较纯2D基线提升了42%的成功率。
原文摘要 · Abstract (English)
Learned world models hold significant potential as neural simulators for robotic manipulation. However, prevalent 2D video-based models inherently lack the spatial and kinematic reasoning crucial for physical interactions. We introduce RoDyn, a novel Robot-Dynamic 2.5D World Model that formulates environmental dynamics within a highly efficient, geometry-aware latent space. Through the proposed Robot-Dynamic Tokenizer, we explicitly couple semantic visual appearances with spatial and agent-centric priors via an RGB-dominated cross-attention mechanism and dynamic mask guidance. Furthermore, by injecting these mask priors directly into sequence transitions, our Mask-guided Autoregressive architecture drives the model to focus on active robot-object interaction regions. Extensive experiments demonstrate that RoDyn establishes SOTA generation fidelity across large-scale datasets. Crucially, it translates these predictive capabilities into substantial downstream gains, accelerating model-based reinforcement learning and achieving a 42\% improvement in real-world imitation learning success rates over pure 2D baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。