arXiv:2411.18135cs.CV2024-11被引 2

用参考图引导生成,解决文本转3D模型模糊低质问题

ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation using Reference Image Prompts

  • 引入参考图像指导的分数蒸馏损失,选择特定生成模式
  • 在T3Bench上实现更高质量、更稳定且更快的3D生成
  • 兼容IP-Adapter,可同时降低估计方差提升效率

基于分数蒸馏采样(SDS)的方法显著推动了文本到3D生成的发展,但现有方法生成的3D模型常出现过平滑和质量低下问题。这源于当前方法的模式搜寻行为:用于更新模型的分数在多个模式间震荡,导致优化不稳定、输出质量下降。为此,我们提出一种新的图像提示分数蒸馏损失(ISD),利用参考图像将文本到3D的优化导向特定模式。ISD可通过IP-Adapter实现,该轻量级适配器可将图像提示能力集成至文生图扩散模型中,作为模式选择模块。其变体在无参考图像时,可作为高效控制变量,降低分数估计方差,从而提升输出质量和优化稳定性。实验表明,ISD在T3Bench基准测试中持续生成视觉连贯、高质量的3D内容,并显著提升优化速度,结果通过定性与定量评估验证。

原文摘要 · Abstract (English)

Existing Score Distillation Sampling (SDS)-based methods have driven significant progress in text-to-3D generation. However, 3D models produced by SDS-based methods tend to exhibit over-smoothing and low-quality outputs. These issues arise from the mode-seeking behavior of current methods, where the scores used to update the model oscillate between multiple modes, resulting in unstable optimization and diminished output quality. To address this problem, we introduce a novel image prompt score distillation loss named ISD, which employs a reference image to direct text-to-3D optimization toward a specific mode. Our ISD loss can be implemented by using IP-Adapter, a lightweight adapter for integrating image prompt capability to a text-to-image diffusion model, as a mode-selection module. A variant of this adapter, when not being prompted by a reference image, can serve as an efficient control variate to reduce variance in score estimates, thereby enhancing both output quality and optimization stability. Our experiments demonstrate that the ISD loss consistently achieves visually coherent, high-quality outputs and improves optimization speed compared to prior text-to-3D methods, as demonstrated through both qualitative and quantitative evaluations on the T3Bench benchmark suite.

3D生成扩散模型图像引导模式控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。