用3D高斯点云实现可交互的关节物体建模与物理推理
ArtGS:3D Gaussian Splatting for Interactive Visual-Physical Modeling and Manipulation of Articulated Objects
- 基于多视角重建+视觉语言模型,自动识别关节结构
- 动态可微渲染优化关节参数,提升运动一致性与操作成功率
- 适合机器人抓取、虚拟仿真等需要物理理解的场景
关节物体操作在机器人领域仍具挑战,源于复杂的运动约束和现有方法物理推理能力有限。本文提出ArtGS,通过扩展3D高斯点云(3DGS)框架,引入视觉-物理联合建模,实现对关节物体的理解与交互。该方法从多视角RGB-D数据重建开始,利用视觉语言模型(VLM)提取语义与结构信息,特别是关节骨骼。通过动态可微的3DGS渲染,优化关节参数,确保运动符合物理约束,并增强操作策略。借助动态高斯点云、跨体适配性及闭环优化,实现了高效、可扩展、通用的关节物体建模与操作。仿真与真实环境实验表明,ArtGS在多种关节物体上显著优于以往方法,关节估计准确率与操作成功率均有提升。更多图片视频见项目主页:https://sites.google.com/view/artgs/home
原文摘要 · Abstract (English)
Articulated object manipulation remains a critical challenge in robotics due to the complex kinematic constraints and the limited physical reasoning of existing methods. In this work, we introduce ArtGS, a novel framework that extends 3D Gaussian Splatting (3DGS) by integrating visual-physical modeling for articulated object understanding and interaction. ArtGS begins with multi-view RGB-D reconstruction, followed by reasoning with a vision-language model (VLM) to extract semantic and structural information, particularly the articulated bones. Through dynamic, differentiable 3DGS-based rendering, ArtGS optimizes the parameters of the articulated bones, ensuring physically consistent motion constraints and enhancing the manipulation policy. By leveraging dynamic Gaussian splatting, cross-embodiment adaptability, and closed-loop optimization, ArtGS establishes a new framework for efficient, scalable, and generalizable articulated object modeling and manipulation. Experiments conducted in both simulation and real-world environments demonstrate that ArtGS significantly outperforms previous methods in joint estimation accuracy and manipulation success rates across a variety of articulated objects. Additional images and videos are available on the project website: https://sites.google.com/view/artgs/home
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。