arXiv:2512.03566cs.CVcs.MM2025-12中稿 · ACM MM Asia2026被引 1

用文本生成可动3D物体,通过三阶段扩散模型实现

GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models

  • 分三步生成:先粗略生成点云,再用超图优化部件结构,最后生成关节连接
  • 在PartNet-Mobility数据集上优于现有方法,能准确还原文本描述的可动结构
  • 适合做3D内容生成、智能设计辅助的开发者和研究者

可动物体生成近年来取得进展,但现有模型难以根据文本提示进行控制。为弥合文本描述与3D可动物体表示之间的差距,我们提出GAOT,一种三阶段框架,通过扩散模型与超图学习,从文本提示生成可动物体。首先,微调点云生成模型以根据文本生成粗略物体表示;其次,利用可动物体与图结构的内在关联,设计基于超图的学习方法,将物体部件表示为图顶点;最后,借助扩散模型,基于部件信息生成关节(以图边表示)。在PartNet-Mobility数据集上的大量定性与定量实验表明,该方法有效,性能优于先前方法。

原文摘要 · Abstract (English)

Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object representations, we propose GAOT, a three-phase framework that generates articulated objects from text prompts, leveraging diffusion models and hypergraph learning in a three-step process. First, we fine-tune a point cloud generation model to produce a coarse representation of objects from text prompts. Given the inherent connection between articulated objects and graph structures, we design a hypergraph-based learning method to refine these coarse representations, representing object parts as graph vertices. Finally, leveraging a diffusion model, the joints of articulated objects-represented as graph edges-are generated based on the object parts. Extensive qualitative and quantitative experiments on the PartNet-Mobility dataset demonstrate the effectiveness of our approach, achieving superior performance over previous methods.

3D生成扩散模型文本生成可动物体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。