通过注意力头重要性向量,让文生图模型更懂人类视觉概念。
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
- 构建注意力头重要性向量,量化每个头对视觉概念的贡献。
- 能减少多义词误生成,成功修改五类复杂图像属性。
- 适合想精细控制生成过程的研究者与开发者使用。
近期文生图扩散模型利用交叉注意力层,有效提升了多种视觉生成任务的表现。然而,对交叉注意力层的理解仍较有限。本文提出一种机制可解释性方法,通过构建与人类指定视觉概念对齐的注意力头相关性向量(HRV)。每个HRV长度等于所有交叉注意力头数量,元素表示对应头对特定概念的重要性。为验证其可解释性,我们设计有序削弱分析并证明其有效性。进一步提出概念增强与调整方法,并应用于三项视觉生成任务。结果表明,HRV可减少多义词在图像生成中的误解释,成功修改五类挑战性属性,缓解多概念生成中的灾难性忽略问题。本研究深化了对交叉注意力层的理解,提出了在头级别精细控制的新方法。
原文摘要 · Abstract (English)
Recent text-to-image diffusion models leverage cross-attention layers, which have been effectively utilized to enhance a range of visual generative tasks. However, our understanding of cross-attention layers remains somewhat limited. In this study, we introduce a mechanistic interpretability approach for diffusion models by constructing Head Relevance Vectors (HRVs) that align with human-specified visual concepts. An HRV for a given visual concept has a length equal to the total number of cross-attention heads, with each element representing the importance of the corresponding head for the given visual concept. To validate HRVs as interpretable features, we develop an ordered weakening analysis that demonstrates their effectiveness. Furthermore, we propose concept strengthening and concept adjusting methods and apply them to enhance three visual generative tasks. Our results show that HRVs can reduce misinterpretations of polysemous words in image generation, successfully modify five challenging attributes in image editing, and mitigate catastrophic neglect in multi-concept generation. Overall, our work provides an advancement in understanding cross-attention layers and introduces new approaches for fine-controlling these layers at the head level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。