让3D生成更懂语义,按部件精准匹配文本描述
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
- 用双通道潜变量同时建模部件几何与外观
- 通过语言推导部件间关系,提升生成一致性
- 适合需要语义对齐的3D内容创作场景
人类认知和推理依赖于对3D物体由有意义部件组成的理解。然而,多数文本到3D方法忽略部件的语义与功能结构。尽管近期部分感知方法引入了分解机制,仍以几何为主,缺乏语义支撑,无法捕捉部件与文本描述的对应关系及其相互依赖。本文提出DreamPartGen,一种语义锚定、部件感知的文本到3D生成框架。该框架引入双重部件潜变量(DPLs)联合建模每个部件的几何与外观,并设计关系语义潜变量(RSLs)从语言中提取部件间依赖关系。通过同步协同去噪过程,强制几何与语义的一致性,实现连贯、可解释且与文本对齐的3D合成。在多个基准测试中,DreamPartGen在几何保真度和文本-形状对齐上达到领先性能。
原文摘要 · Abstract (English)
Understanding and generating 3D objects as compositions of meaningful parts is fundamental to human perception and reasoning. However, most text-to-3D methods overlook the semantic and functional structure of parts. While recent part-aware approaches introduce decomposition, they remain largely geometry-focused, lacking semantic grounding and failing to model how parts align with textual descriptions or their inter-part relations. We propose DreamPartGen, a framework for semantically grounded, part-aware text-to-3D generation. DreamPartGen introduces Duplex Part Latents (DPLs) that jointly model each part's geometry and appearance, and Relational Semantic Latents (RSLs) that capture inter-part dependencies derived from language. A synchronized co-denoising process enforces mutual geometric and semantic consistency, enabling coherent, interpretable, and text-aligned 3D synthesis. Across multiple benchmarks, DreamPartGen delivers state-of-the-art performance in geometric fidelity and text-shape alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。