arXiv:2412.02287cs.CV2024-12ICCV被引 2

用注意力与CLIP引导,解决3D生成视角不一致问题

Improving Viewpoint Consistency in 3D Generation via Structure Feature and CLIP Guidance

  • 通过自适应控制交叉注意力图增强目标视角
  • 利用CLIP视图-文本相似度过滤错误视角,提升一致性
  • 无需微调,可直接嵌入现有3D生成框架

尽管文本到3D生成技术取得进展,当前方法仍常因几何不一致(即'雅努斯问题')而受限。本文识别出该问题根源:扩散模型中的视角生成偏差,导致实际生成视角与优化3D模型所需视角存在显著差距。为此,提出无需微调的注意力与CLIP引导(ACG)机制:通过自适应调控交叉注意力图强化目标视角,利用基于CLIP的视图-文本相似度筛选错误视角,并采用分阶段提示的粗到精优化策略逐步提升3D生成质量。大量实验表明,该方法显著缓解雅努斯问题,且不降低生成速度,可作为现有文本到3D框架的高效即插即用组件。

原文摘要 · Abstract (English)

Despite recent advances in text-to-3D generation techniques, current methods often suffer from geometric inconsistencies, commonly referred to as the Janus Problem. This paper identifies the root cause of the Janus Problem: viewpoint generation bias in diffusion models, which creates a significant gap between the actual generated viewpoint and the expected one required for optimizing the 3D model. To address this issue, we propose a tuning-free approach called the Attention and CLIP Guidance (ACG) mechanism. ACG enhances desired viewpoints by adaptively controlling cross-attention maps, employs CLIP-based view-text similarities to filter out erroneous viewpoints, and uses a coarse-to-fine optimization strategy with staged prompts to progressively refine 3D generation. Extensive experiments demonstrate that our method significantly reduces the Janus Problem without compromising generation speed, establishing ACG as an efficient, plug-and-play component for existing text-to-3D frameworks.

3D生成视角一致CLIP扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。