arXiv:2604.05906cs.CVcs.AI2026-04被引 1

选相关注意力头聚合,提升文本生成图像的可解释性。

Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation

  • 按目标概念筛选关键注意力头进行聚合
  • 比DAAM方法平均交并比更高
  • 能诊断提示词误解,适合可解释性研究

大量关于文本到图像(T2I)生成模型的研究利用交叉注意力图来提升应用性能并解释模型行为。然而,不同注意力头产生的注意力图的特性仍相对未被深入探索。本研究发现,选择与目标概念最相关的注意力头进行聚合,可显著提升视觉可解释性。相较于基于扩散的分割方法DAAM,该方法在平均交并比(mean IoU)上表现更优。我们还发现,最相关注意力头能更准确捕捉概念特异性特征,而最少相关头则表现较差。通过选择性聚合,还能有效诊断提示词的误理解问题。这些结果表明,注意力头选择为提升T2I生成的可解释性与可控性提供了有前景的方向。

原文摘要 · Abstract (English)

Numerous studies on text-to-image (T2I) generative models have utilized cross-attention maps to boost application performance and interpret model behavior. However, the distinct characteristics of attention maps from different attention heads remain relatively underexplored. In this study, we show that selectively aggregating cross-attention maps from heads most relevant to a target concept can improve visual interpretability. Compared to the diffusion-based segmentation method DAAM, our approach achieves higher mean IoU scores. We also find that the most relevant heads capture concept-specific features more accurately than the least relevant ones, and that selective aggregation helps diagnose prompt misinterpretations. These findings suggest that attention head selection offers a promising direction for improving the interpretability and controllability of T2I generation.

文本生成图像注意力机制可解释性扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。