arXiv:2410.21299cs.CV2024-10TPAMI被引 10

解决文本+图像提示生成3D模型时质量差的问题

TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt

  • 用确定性噪声替代随机噪声,改进Score Distillation Sampling
  • 融合视觉提示与文本信息,生成高质量定制化3D模型
  • 适合需要多条件控制的3D内容创作用户

近年来,生成模型在文本到3D生成方面取得显著进展。许多方法依赖于得分蒸馏采样(SDS)技术,但SDS难以支持多条件输入(如文本和视觉提示)的定制化生成任务。我们分析发现,问题源于SDS中的差异项及优化过程中的随机噪声,导致蒸馏过程中偏离目标模式。为此,我们提出新的分类器得分匹配(CSM)算法,移除SDS中的差异项,并采用确定性噪声添加方式以减少优化过程中的噪声,有效克服了SDS在定制化生成中的低质量缺陷。基于CSM,我们引入视觉提示融合机制与采样引导技术,构建视觉提示CSM(VPCSM)算法;同时设计语义-几何校准(SGC)模块,提升文本信息整合能力。我们提出TV-3DG框架,大量实验表明其能实现稳定、高质量的定制化3D生成。

原文摘要 · Abstract (English)

In recent years, advancements in generative models have significantly expanded the capabilities of text-to-3D generation. Many approaches rely on Score Distillation Sampling (SDS) technology. However, SDS struggles to accommodate multi-condition inputs, such as text and visual prompts, in customized generation tasks. To explore the core reasons, we decompose SDS into a difference term and a classifier-free guidance term. Our analysis identifies the core issue as arising from the difference term and the random noise addition during the optimization process, both contributing to deviations from the target mode during distillation. To address this, we propose a novel algorithm, Classifier Score Matching (CSM), which removes the difference term in SDS and uses a deterministic noise addition process to reduce noise during optimization, effectively overcoming the low-quality limitations of SDS in our customized generation framework. Based on CSM, we integrate visual prompt information with an attention fusion mechanism and sampling guidance techniques, forming the Visual Prompt CSM (VPCSM) algorithm. Furthermore, we introduce a Semantic-Geometry Calibration (SGC) module to enhance quality through improved textual information integration. We present our approach as TV-3DG, with extensive experiments demonstrating its capability to achieve stable, high-quality, customized 3D generation. Project page: \url{https://yjhboy.github.io/TV-3DG}

3D生成文本生成视觉提示扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。