从稀疏视频中学习物理驱动的世界模型,精准捕捉复杂接触交互。
ContactGaussian-WM: Learning Physics-Grounded World Model from Videos
- 用统一高斯表示视觉与碰撞几何,实现物理感知建模。
- 在稀疏视频下学习复杂动态场景,性能超越现有最先进方法。
- 适用于机器人规划、实时控制等下游任务,具强泛化能力。
构建能理解复杂物理交互的世界模型,对推动机器人规划与仿真至关重要。然而,现有方法在数据稀缺和接触丰富的动态运动条件下常表现不佳。为此,我们提出 ContactGaussian-WM,一种可微分的物理驱动刚体世界模型,能够直接从稀疏且接触密集的视频序列中学习复杂的物理规律。该框架包含两个核心组件:(1) 统一的高斯表示,同时建模视觉外观与碰撞几何;(2) 端到端可微学习框架,通过闭式物理引擎反向传播,从稀疏视觉观测中推断物理属性。大量仿真与真实世界评估表明,ContactGaussian-WM 在学习复杂场景时优于现有最先进方法,展现出强大的泛化能力。此外,我们展示了该框架在数据合成与实时模型预测控制(MPC)中的实用价值。
原文摘要 · Abstract (English)
Developing world models that understand complex physical interactions is essential for advancing robotic planning and simulation.However, existing methods often struggle to accurately model the environment under conditions of data scarcity and complex contact-rich dynamic motion.To address these challenges, we propose ContactGaussian-WM, a differentiable physics-grounded rigid-body world model capable of learning intricate physical laws directly from sparse and contact-rich video sequences.Our framework consists of two core components: (1) a unified Gaussian representation for both visual appearance and collision geometry, and (2) an end-to-end differentiable learning framework that differentiates through a closed-form physics engine to infer physical properties from sparse visual observations.Extensive simulations and real-world evaluations demonstrate that ContactGaussian-WM outperforms state-of-the-art methods in learning complex scenarios, exhibiting robust generalization capabilities.Furthermore, we showcase the practical utility of our framework in downstream applications, including data synthesis and real-time MPC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。