通过动态锐化分数函数,减少扩散模型生成中的幻觉问题。
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
- 仅在已知导致伪影的方向上锐化分数函数,保留有效语义变化。
- 在控制和自然图像数据集上显著降低幻觉率,优于现有基线。
- 可扩展至文生图任务,基于文本描述的细粒度语义差异选择模式。
扩散模型中的幻觉表现为结构不一致的样本,源于学习到的分数函数过度平滑,进而导致数据分布模式间的插值。尽管语义插值常带来样本多样性,但需针对性解决幻觉问题。本文提出动态引导(Dynamic Guidance),仅在预设的、易引发伪影的方向上选择性锐化分数函数,同时保留有效的语义变化。该方法可基于预定义类别或数据分布中形成的语义一致聚类(伪类别)实现。后者使动态引导能自然拓展至文生图任务,其中模式对应文本描述中的细粒度上下文差异。据我们所知,这是首个在生成时直接缓解幻觉而非依赖事后过滤的方法。在控制与自然图像数据集上,动态引导显著减少幻觉,性能远超基线。
原文摘要 · Abstract (English)
Hallucinations in diffusion models are samples with structural inconsistencies that can emerge due to the excessive smoothing of the learned score function, which in turn leads to interpolations between modes of the data distribution. Since semantic interpolations are often desirable and contribute to sample diversity, we believe that a nuanced and targeted solution is required to address diffusion model hallucinations. In this work, we introduce Dynamic Guidance, which mitigates hallucinations by selectively sharpening the score function only along the pre-determined directions known to cause artifacts, while preserving valid semantic variations. This sharpening can be performed using either pre-determined classes or semantically coherent clusters that form pseudo-classes over the data distribution. The latter allows for a principled extension of Dynamic Guidance to text-to-image generation, where we select modes to correspond to fine-grained contextual differences in textual descriptions. To our knowledge, this is the first approach that addresses hallucinations at generation time rather than through post-hoc filtering. Dynamic Guidance substantially reduces hallucinations on both controlled and natural image datasets, significantly outperforming baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。