用视觉语言模型自动构建可交互3D物体,省去人工建模
Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model

- 基于多模态输入生成可编译的交互代码
- 在PartNet-Mobility上成功率从不足12%提升至75%
- 适合想快速生成3D交互资产的开发者和研究者
交互式3D模拟物体在AR/VR、动画和机器人领域至关重要,但传统建模需大量人工。我们提出Articulate-Anything,通过视觉语言模型(VLMs)自动实现多种复杂物体的关节化,支持文本、图像和视频输入。系统利用网格检索机制调用现有3D资产数据集,并采用演员-评论家框架迭代优化关节配置,自我修正错误,生成可在标准3D仿真器中使用的可交互数字孪生体。定性评估显示其能有效处理复杂甚至模糊的物体功能。在标准PartNet-Mobility数据集上的定量实验表明,成功率从8.7%-11.6%大幅提升至75%,达到新基准。我们还展示了从真实视频生成3D资产的能力,用于训练精细操作策略,在仿真中完成超越基础抓取的任务,并成功部署到真实机器人系统。
原文摘要 · Abstract (English)
Interactive 3D simulated objects are crucial in AR/VR, animations, and robotics, driving immersive experiences and advanced automation. However, creating these articulated objects requires extensive human effort and expertise, limiting their broader applications. To overcome this challenge, we present Articulate-Anything, a system that automates the articulation of diverse, complex objects from many input modalities, including text, images, and videos. Articulate-Anything leverages vision-language models (VLMs) to generate code that can be compiled into an interactable digital twin for use in standard 3D simulators. Our system exploits existing 3D asset datasets via a mesh retrieval mechanism, along with an actor-critic system that iteratively proposes, evaluates, and refines solutions for articulating the objects, self-correcting errors to achieve a robust outcome. Qualitative evaluations demonstrate Articulate-Anything's capability to articulate complex and even ambiguous object affordances by leveraging rich grounded inputs. In extensive quantitative experiments on the standard PartNet-Mobility dataset, Articulate-Anything substantially outperforms prior work, increasing the success rate from 8.7-11.6% to 75% and setting a new bar for state-of-the-art performance. We further showcase the utility of our system by generating 3D assets from in-the-wild video inputs, which are then used to train robotic policies for fine-grained manipulation tasks in simulation that go beyond basic pick and place. These policies are then transferred to a real robotic system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。