用运动约束指导3D物体分解,实现跨类别高泛化重建。
Instruct-Particulate: Scaling Feed-Forward 3D Object Articulation with Kinematic Control

- 输入3D网格+运动结构描述,预测部件分割与关节参数
- 在15万+异构数据上训练,支持任意输入网格的自动解析
- 适合游戏/机器人领域需快速生成可动3D资产的场景
重建可动3D物体对动画、游戏和机器人仿真至关重要。现有神经网络虽能估计3D物体的结构,但受限于标注数据稀缺,泛化能力不足。为此,我们提出Instruct-Particulate,该模型接收3D网格及目标运动学规格(包括部件描述、连接关系、关节类型和可选点提示),预测对应的运动部件分割与关节运动参数。运动学规格明确任务边界,使模型可适配不同粒度的标注,从而利用更丰富的异构训练数据。测试时,运动学规格可通过大规模视觉-语言模型自动生成,使模型适用于任意输入网格。为实现规模化训练,我们构建了一个包含超过15万件可动3D物体的异构数据集,通过视觉-语言模型对其他3D模型(整体或已分解)进行部分标注扩展了现有公开数据集。实验表明,该模型在跨类别和生成网格上均具有更强泛化能力,支持通过图像到3D模型实现真实世界图像中可动资产的重建。
原文摘要 · Abstract (English)
Reconstructing articulated 3D objects is important for animation, gaming, and robotic simulations. Recent neural networks can estimate the articulated structure of 3D objects, but their generalization remains limited by the scarcity of annotated data for this task. To address this gap, we introduce Instruct-Particulate, a model that takes a 3D mesh together with a target kinematic specification, including part descriptions, connectivity, joint types, and optional point prompts, and predicts the corresponding kinematic part segmentation and joint motion parameters. The kinematic specification disambiguates the task and allows the model to target annotations of different granularity, thereby making it possible to use more abundant heterogeneous training data. At test time, the kinematic specification can be obtained automatically from large-scale vision-language models, so the model can be applied to any input mesh. To train our model at scale, we construct a heterogeneous dataset of more than 150,000 articulated 3D objects, extending existing publicly available collections with data obtained by partially labelling other 3D models (monolithic or already decomposed into parts) with kinematic labels by means of vision-language models. Experiments show that our model generalizes better across categories and to AI-generated meshes, enabling articulated asset reconstruction from real-world images via image-to-3D models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。