让机器人视频生成更真实,融入物理规律建模3D结构与运动。
RoboScape: Physics-informed Embodied World Model
- 联合训练视觉生成与物理知识,统一建模视频与物理规律。
- 通过时序深度预测和关键点动力学学习,提升三维一致性与运动合理性。
- 适合研究机器人仿真、强化学习策略训练的学者使用。
世界模型已成为具身智能的关键工具,可生成逼真的机器人视频并缓解数据稀缺问题。然而,现有具身世界模型对物理规律感知不足,尤其在3D几何与运动动力学建模方面表现有限,导致接触密集场景下视频生成不真实。本文提出RoboScape,一个统一的物理感知世界模型,在集成框架中联合学习RGB视频生成与物理知识。引入两项关键的物理感知联合训练任务:时序深度预测以增强视频渲染中的3D几何一致性;关键点动力学学习隐式编码物体形状与材质等物理属性,同时提升复杂运动建模能力。大量实验表明,RoboScape在多种机器人场景下生成的视频具有更高的视觉保真度与物理合理性。进一步通过下游应用验证其实用性,包括基于生成数据的机器人策略训练与策略评估。本工作为构建高效物理感知世界模型提供了新思路,推动具身智能研究发展。代码已开源:https://github.com/tsinghua-fib-lab/RoboScape。
原文摘要 · Abstract (English)
World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models exhibit limited physical awareness, particularly in modeling 3D geometry and motion dynamics, resulting in unrealistic video generation for contact-rich robotic scenarios. In this paper, we present RoboScape, a unified physics-informed world model that jointly learns RGB video generation and physics knowledge within an integrated framework. We introduce two key physics-informed joint training tasks: temporal depth prediction that enhances 3D geometric consistency in video rendering, and keypoint dynamics learning that implicitly encodes physical properties (e.g., object shape and material characteristics) while improving complex motion modeling. Extensive experiments demonstrate that RoboScape generates videos with superior visual fidelity and physical plausibility across diverse robotic scenarios. We further validate its practical utility through downstream applications including robotic policy training with generated data and policy evaluation. Our work provides new insights for building efficient physics-informed world models to advance embodied intelligence research. The code is available at: https://github.com/tsinghua-fib-lab/RoboScape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。