arXiv:2505.20460cs.CV2025-05NeurIPS被引 21

用两张图生成可动3D物体,提升结构合理性与泛化能力

DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data

  • 双图像输入捕捉静止与运动状态,指导关节关系预测
  • 生成的物体在静止和运动状态上均优于现有方法
  • 构建大规模新数据集PM-X,支持复杂可动物体建模

我们提出DIPO,一种从一对图像中可控生成可动3D物体的新框架:一张为物体静止状态,另一张为运动状态。相比单图方法,双图输入仅小幅增加数据收集成本,但提供了关键的运动信息,有效引导部件间运动关系的预测。我们设计了双图像扩散模型,捕捉图像对间的关联,生成部件布局与关节参数;同时引入基于思维链(CoT)的图推理器,显式推断部件连接关系。为提升复杂可动物体的鲁棒性与泛化能力,我们开发了全自动数据集扩展流程LEGO-Art,扩充PartNet-Mobility数据集。我们还提出了PM-X,一个包含复杂可动3D物体的大规模数据集,附带渲染图像、URDF标注和文本描述。大量实验表明,DIPO在静止和运动状态下均显著优于现有基线,且PM-X数据集进一步增强了对多样、复杂结构的泛化能力。代码与数据集将在发表后公开。

原文摘要 · Abstract (English)

We present DIPO, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach, our dual-image input imposes only a modest overhead for data collection, but at the same time provides important motion information, which is a reliable guide for predicting kinematic relationships between parts. Specifically, we propose a dual-image diffusion model that captures relationships between the image pair to generate part layouts and joint parameters. In addition, we introduce a Chain-of-Thought (CoT) based graph reasoner that explicitly infers part connectivity relationships. To further improve robustness and generalization on complex articulated objects, we develop a fully automated dataset expansion pipeline, name LEGO-Art, that enriches the diversity and complexity of PartNet-Mobility dataset. We propose PM-X, a large-scale dataset of complex articulated 3D objects, accompanied by rendered images, URDF annotations, and textual descriptions. Extensive experiments demonstrate that DIPO significantly outperforms existing baselines in both the resting state and the articulated state, while the proposed PM-X dataset further enhances generalization to diverse and structurally complex articulated objects. Our code and dataset will be released to the community upon publication.

3D生成可动物体扩散模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。