用物理感知世界模型提升机器人推拉操作的泛化能力
PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation
- 基于可微分物理仿真,从少量视觉轨迹中端到端学习3D刚体动力学
- 仅需少样本交互数据即可实现高精度动力学建模,且无需状态估计
- 通过物理随机化生成数字孪生体,有效解决仿真到现实的迁移难题
非抓取式操作(如推、戳)是机器人基础技能,但因摩擦与恢复力等复杂物理交互而难以学习。为实现稳健策略学习与泛化,本文提出物理感知世界模型PIN-WM,通过可微分物理仿真,仅用少量无任务依赖的物理交互轨迹,即可端到端识别3D刚体动力学系统。PIN-WM利用高斯点云的观测损失进行训练,无需状态估计。为弥合仿真到现实的差距,通过物理感知随机化生成一组数字孪生体,扰动物理与渲染参数以生成多样化且有意义的模型变体。在仿真和真实场景的广泛测试中,结合物理感知数字孪生体的PIN-WM,在仿真到现实迁移中表现优于现有Real2Sim2Real方法。
原文摘要 · Abstract (English)
While non-prehensile manipulation (e.g., controlled pushing/poking) constitutes a foundational robotic skill, its learning remains challenging due to the high sensitivity to complex physical interactions involving friction and restitution. To achieve robust policy learning and generalization, we opt to learn a world model of the 3D rigid body dynamics involved in non-prehensile manipulations and use it for model-based reinforcement learning. We propose PIN-WM, a Physics-INformed World Model that enables efficient end-to-end identification of a 3D rigid body dynamical system from visual observations. Adopting differentiable physics simulation, PIN-WM can be learned with only few-shot and task-agnostic physical interaction trajectories. Further, PIN-WM is learned with observational loss induced by Gaussian Splatting without needing state estimation. To bridge Sim2Real gaps, we turn the learned PIN-WM into a group of Digital Cousins via physics-aware randomizations which perturb physics and rendering parameters to generate diverse and meaningful variations of the PIN-WM. Extensive evaluations on both simulation and real-world tests demonstrate that PIN-WM, enhanced with physics-aware digital cousins, facilitates learning robust non-prehensile manipulation skills with Sim2Real transfer, surpassing the Real2Sim2Real state-of-the-arts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。