用3D多模态大模型自动建模可动物体,提升仿真精度。
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
- 基于点云与文本的自回归预测,联合优化分割与运动参数。
- 几何分割mIoU提升17%,运动参数误差降低29%,物理可执行性提高50%。
- 无需人工干预,对未见物体也表现良好,适合机器人仿真应用。
构建可动物体的精确数字孪生体对机器人模拟训练和具身智能世界建模至关重要,但传统方法依赖繁琐的手动建模或多阶段流程。本文提出基于3D多模态大语言模型(MLLM)的端到端自动重建框架URDF-Anything。该方法利用点云与文本的多模态输入,通过自回归预测框架联合优化几何分割与运动学参数预测,并设计专用的$[SEG]$标记机制,直接与点云特征交互,实现细粒度部件级分割并保持与运动参数预测的一致性。在模拟与真实数据集上的实验表明,该方法在几何分割(mIoU提升17%)、运动学参数预测(平均误差降低29%)及物理可执行性(超越基线50%)方面显著优于现有方法。尤其具备出色的泛化能力,对训练集外物体仍表现优异。本工作为机器人模拟中的数字孪生构建提供了高效解决方案,显著增强从仿真到现实的迁移能力。
原文摘要 · Abstract (English)
Constructing accurate digital twins of articulated objects is essential for robotic simulation training and embodied AI world model building, yet historically requires painstaking manual modeling or multi-stage pipelines. In this work, we propose \textbf{URDF-Anything}, an end-to-end automatic reconstruction framework based on a 3D multimodal large language model (MLLM). URDF-Anything utilizes an autoregressive prediction framework based on point-cloud and text multimodal input to jointly optimize geometric segmentation and kinematic parameter prediction. It implements a specialized $[SEG]$ token mechanism that interacts directly with point cloud features, enabling fine-grained part-level segmentation while maintaining consistency with the kinematic parameter predictions. Experiments on both simulated and real-world datasets demonstrate that our method significantly outperforms existing approaches regarding geometric segmentation (mIoU 17\% improvement), kinematic parameter prediction (average error reduction of 29\%), and physical executability (surpassing baselines by 50\%). Notably, our method exhibits excellent generalization ability, performing well even on objects outside the training set. This work provides an efficient solution for constructing digital twins for robotic simulation, significantly enhancing the sim-to-real transfer capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。