arXiv:2603.01142cs.CV2026-03被引 4

用3D大模型一键生成可动3D物体,自动拆解零件并设计关节。

ArtLLM: Generating Articulated Assets via 3D LLM

  • 基于3D多模态大模型,从完整网格自动生成零件与关节结构。
  • 在PartNet-Mobility上零件布局和关节预测准确率显著超越现有方法。
  • 适合游戏、机器人仿真等需要大量可动3D资产的场景。

为游戏、机器人和仿真构建交互式数字环境,依赖于由部件几何形状和运动结构共同决定功能的可动3D物体。然而,现有方法存在根本局限:基于优化的重建方法需逐个物体进行关节拟合,速度慢且仅支持简单单关节对象;基于检索的方法则从固定库中拼接部件,导致几何重复且泛化能力差。为此,我们提出ArtLLM,一种直接从完整3D网格生成高质量可动资产的新框架。其核心是一个在大规模人工标注数据集上训练的3D多模态大语言模型,该数据集来自现有标注数据集与程序生成对象的结合。与以往工作不同,ArtLLM能自回归地预测可变数量的零件与关节,并统一推断其运动结构,输入为物体点云。该具运动感知的布局随后作为条件,驱动3D生成模型合成高保真部件几何。在PartNet-Mobility数据集上的实验表明,ArtLLM在零件布局准确性和关节预测方面显著优于当前最佳方法,并对真实世界物体具有强泛化能力。最后,我们展示了其在构建数字孪生中的应用,凸显其在可扩展机器人学习中的潜力。

原文摘要 · Abstract (English)

Creating interactive digital environments for gaming, robotics, and simulation relies on articulated 3D objects whose functionality emerges from their part geometry and kinematic structure. However, existing approaches remain fundamentally limited: optimization-based reconstruction methods require slow, per-object joint fitting and typically handle only simple, single-joint objects, while retrieval-based methods assemble parts from a fixed library, leading to repetitive geometry and poor generalization. To address these challenges, we introduce ArtLLM, a novel framework for generating high-quality articulated assets directly from complete 3D meshes. At its core is a 3D multimodal large language model trained on a large-scale articulation dataset curated from both existing articulation datasets and procedurally generated objects. Unlike prior work, ArtLLM autoregressively predicts a variable number of parts and joints, inferring their kinematic structure in a unified manner from the object's point cloud. This articulation-aware layout then conditions a 3D generative model to synthesize high-fidelity part geometries. Experiments on the PartNet-Mobility dataset show that ArtLLM significantly outperforms state-of-the-art methods in both part layout accuracy and joint prediction, while generalizing robustly to real-world objects. Finally, we demonstrate its utility in constructing digital twins, highlighting its potential for scalable robot learning.

3D生成可动资产大模型数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。