arXiv:2410.09009cs.CV2024-10被引 11

用语义引导的扩散采样提升文本生成3D的细节与组合能力

Semantic Score Distillation Sampling for Compositional Text-to-3D Generation

  • 通过语义嵌入构建视图一致的语义图,实现区域化精细化优化
  • 在Objaverse和ShapeNet数据集上生成复杂3D场景质量显著提升
  • 适合需要精确控制物体布局与结构的3D内容创作者

从文本描述生成高质量3D资产仍是计算机图形与视觉领域的关键挑战。由于3D数据稀缺,当前先进方法利用预训练2D扩散模型,并通过分数蒸馏采样(SDS)进行优化。尽管取得进展,生成包含多个物体或复杂交互的复杂3D场景仍困难重重。现有布局引导方法多依赖粗粒度框或布局提示,缺乏细粒度控制能力。为此,本文提出一种新型SDS方法——语义分数蒸馏采样(SemanticSDS),通过引入跨视角保持一致且能清晰区分不同物体与部件的语义嵌入,构建语义图,指导区域特异性SDS过程,实现精准优化与组合生成。该方法充分释放现有预训练扩散模型的组合潜力,在复杂物体与场景生成上达到领先水平。实验表明,SemanticSDS在Objaverse与ShapeNet数据集上均显著优于基线方法。

原文摘要 · Abstract (English)

Generating high-quality 3D assets from textual descriptions remains a pivotal challenge in computer graphics and vision research. Due to the scarcity of 3D data, state-of-the-art approaches utilize pre-trained 2D diffusion priors, optimized through Score Distillation Sampling (SDS). Despite progress, crafting complex 3D scenes featuring multiple objects or intricate interactions is still difficult. To tackle this, recent methods have incorporated box or layout guidance. However, these layout-guided compositional methods often struggle to provide fine-grained control, as they are generally coarse and lack expressiveness. To overcome these challenges, we introduce a novel SDS approach, Semantic Score Distillation Sampling (SemanticSDS), designed to effectively improve the expressiveness and accuracy of compositional text-to-3D generation. Our approach integrates new semantic embeddings that maintain consistency across different rendering views and clearly differentiate between various objects and parts. These embeddings are transformed into a semantic map, which directs a region-specific SDS process, enabling precise optimization and compositional generation. By leveraging explicit semantic guidance, our method unlocks the compositional capabilities of existing pre-trained diffusion models, thereby achieving superior quality in 3D content generation, particularly for complex objects and scenes. Experimental results demonstrate that our SemanticSDS framework is highly effective for generating state-of-the-art complex 3D content. Code: https://github.com/YangLing0818/SemanticSDS-3D

文本生成3D扩散模型语义控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。