arXiv:2411.17472cs.CVcs.LG2024-11被引 4

用贝叶斯理论改进文生图模型的注意力机制,让修饰词更准地对应名词。

Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory

  • 设计可定制先验控制注意力分布,强化对象分离和修饰词绑定
  • 在多个基准上达到当前最优性能,显著提升复杂提示的生成质量
  • 适合关注模型可解释性与生成准确性的研究者与开发者

文生图扩散模型通过文本提示生成高保真、多样且视觉逼真的图像,实现了生成建模的突破。然而,现有模型在涉及多个对象和属性的复杂提示下仍存在修饰词与名词错配或遗漏元素的问题。尽管基于注意力的方法改善了对象包含和语言绑定,但仍面临属性错绑及泛化能力不足的挑战。本文利用PAC-Bayes框架,提出一种贝叶斯方法,在注意力分布上设计定制先验,以强制实现对象间差异、修饰词与对应名词对齐、对无关词元的最小注意力以及更好的泛化正则化。该方法将注意力机制视为可解释组件,实现细粒度控制并提升属性-对象对齐效果。我们在标准基准上验证了方法的有效性,多指标均达当前最优水平。通过在去噪过程中融入定制先验,该方法提升了图像质量,解决了文生图模型长期存在的对齐难题,为更可靠、可解释的生成模型铺平道路。

原文摘要 · Abstract (English)

Text-to-image (T2I) diffusion models have revolutionized generative modeling by producing high-fidelity, diverse, and visually realistic images from textual prompts. Despite these advances, existing models struggle with complex prompts involving multiple objects and attributes, often misaligning modifiers with their corresponding nouns or neglecting certain elements. Recent attention-based methods have improved object inclusion and linguistic binding, but still face challenges such as attribute misbinding and a lack of robust generalization guarantees. Leveraging the PAC-Bayes framework, we propose a Bayesian approach that designs custom priors over attention distributions to enforce desirable properties, including divergence between objects, alignment between modifiers and their corresponding nouns, minimal attention to irrelevant tokens, and regularization for better generalization. Our approach treats the attention mechanism as an interpretable component, enabling fine-grained control and improved attribute-object alignment. We demonstrate the effectiveness of our method on standard benchmarks, achieving state-of-the-art results across multiple metrics. By integrating custom priors into the denoising process, our method enhances image quality and addresses long-standing challenges in T2I diffusion models, paving the way for more reliable and interpretable generative models.

文生图扩散模型注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。