用三维重建+视觉模型+物理仿真,让机器人在真实环境里自主规划动作。
Scan, Materialize, Simulate: A Generalizable Framework for Physically Grounded Robot Planning
- 融合3D高斯泼溅、视觉大模型与物理仿真,实现场景理解与动作预测
- 在台球和无人机着陆任务中均表现稳定,跨仿真到真实环境泛化能力强
- 无需重新学习物理规律,适合需要灵活应对复杂场景的机器人应用
自主机器人必须推理其行为的物理后果,才能在非结构化的真实环境中有效运行。我们提出扫描、构建、模拟(SMS)框架,结合3D高斯泼溅实现精准场景重建,利用视觉基础模型进行语义分割,通过视觉语言模型推断材料属性,并借助物理仿真可靠预测动作结果。通过整合这些组件,SMS实现了无需重学基础物理动态的可泛化物理推理与以物体为中心的规划。我们在一个类台球操作任务和一个具有挑战性的四旋翼无人机着陆场景中进行了实证验证,展示了在仿真域迁移和真实世界实验中的稳健性能。结果表明,将可微渲染、基础模型与基于物理的仿真相结合,可在多种环境下实现物理感知的机器人规划。
原文摘要 · Abstract (English)
Autonomous robots must reason about the physical consequences of their actions to operate effectively in unstructured, real-world environments. We present Scan, Materialize, Simulate (SMS), a unified framework that combines 3D Gaussian Splatting for accurate scene reconstruction, visual foundation models for semantic segmentation, vision-language models for material property inference, and physics simulation for reliable prediction of action outcomes. By integrating these components, SMS enables generalizable physical reasoning and object-centric planning without the need to re-learn foundational physical dynamics. We empirically validate SMS in a billiards-inspired manipulation task and a challenging quadrotor landing scenario, demonstrating robust performance on both simulated domain transfer and real-world experiments. Our results highlight the potential of bridging differentiable rendering for scene reconstruction, foundation models for semantic understanding, and physics-based simulation to achieve physically grounded robot planning across diverse settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。