arXiv:2607.01938cs.ROcs.AI2026-07

用物理规律建模动态物体,让机器人更准地抓动移动目标。

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

论文配图:PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
图 1 · 摘自论文原文
  • 结合物理规律的3D高斯世界模型,实时预测物体运动轨迹。
  • 在16项任务中成功率超越现有方法,真实机器人实验也表现优异。
  • 适合研究物理感知、机器人操控与动态环境建模的开发者。

在非结构化3D环境中操纵快速移动的目标,对具身智能仍具挑战。现有视觉-语言-动作模型与世界模型难以实现准确的3D几何建模和物理合理的未来预测。我们提出PhysMani,将物理原理驱动的3D高斯世界模型与前瞻意识的动作策略模型相结合。世界模型通过在线优化学习无散度的高斯速度场,实现快速且物理可信的未来动态预测。策略模型通过可学习的标记交叉注意力模块融合预测的3D场景未来动态。我们构建了PhysMani-Bench,一个包含16个任务的动态操作基准,结果表明该方法在仿真和真实机器人实验中均显著优于强基线。

原文摘要 · Abstract (English)

Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action models and world models struggle with accurate 3D geometry and physically meaningful forecasting. We propose PhysMani, a framework that couples a physics-principled 3D Gaussian world model with a future-aware action policy model. The world model learns a divergence-free Gaussian velocity field via online optimization for fast and physically grounded future dynamics prediction. The policy model integrates the predicted 3D scene future dynamics through a learnable token based cross-attention module. We introduce PhysMani-Bench, a dynamic manipulation benchmark with 16 tasks, and demonstrate a superior success rate over strong baselines in both simulation and real-world robot experiments.

物理建模3D世界模型机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。