将任意3D模型自动转为可动结构,支持开放词汇建模。
Articulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling
- 用视觉语言模型和提示技术解析语义,分割部件并构建功能关节。
- 生成工具、玩具、机械装置等多样可动3D模型,覆盖范围远超现有数据集。
- 生成资产可用于仿真训练机器人操作技能,并迁移到真实机械臂。
3D可动物体建模长期面临挑战,需同时捕捉精确几何形状与语义清晰、空间精准的部件与关节结构。现有方法依赖有限手选类别(如柜子、抽屉)的训练数据,难以在开放词汇场景中泛化。为此,我们提出 Articulate Anymesh,一个自动化框架,可将任意刚性3D网格转化为其可动对应物。给定3D网格,该框架利用先进的视觉-语言模型与视觉提示技术提取语义信息,实现部件分割与功能关节构建。实验表明,Articulate Anymesh 能生成大规模、高质量的3D可动物体,涵盖工具、玩具、机械装置与车辆,显著扩展了现有3D可动物体数据集的覆盖范围。此外,生成资产可促进仿真中新型可动物体操作技能的学习,并成功迁移到真实机器人系统。项目主页:https://articulate-anymesh.github.io。
原文摘要 · Abstract (English)
3D articulated objects modeling has long been a challenging problem, since it requires to capture both accurate surface geometries and semantically meaningful and spatially precise structures, parts, and joints. Existing methods heavily depend on training data from a limited set of handcrafted articulated object categories (e.g., cabinets and drawers), which restricts their ability to model a wide range of articulated objects in an open-vocabulary context. To address these limitations, we propose Articulate Anymesh, an automated framework that is able to convert any rigid 3D mesh into its articulated counterpart in an open-vocabulary manner. Given a 3D mesh, our framework utilizes advanced Vision-Language Models and visual prompting techniques to extract semantic information, allowing for both the segmentation of object parts and the construction of functional joints. Our experiments show that Articulate Anymesh can generate large-scale, high-quality 3D articulated objects, including tools, toys, mechanical devices, and vehicles, significantly expanding the coverage of existing 3D articulated object datasets. Additionally, we show that these generated assets can facilitate the acquisition of new articulated object manipulation skills in simulation, which can then be transferred to a real robotic system. Our Github website is https://articulate-anymesh.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。