arXiv:2605.09313cs.CV2026-05被引 1

发现扩散模型中的注意力焦点不影响图像对齐,但抑制它们会引发特定感知变化。

Attention Sinks in Diffusion Transformers: A Causal Analysis

论文配图:Attention Sinks in Diffusion Transformers: A Causal Analysis
图 1 · 摘自论文原文
  • 通过动态识别每步关键注意力目标,训练无关地干预得分与值路径。
  • 移除注意力焦点后,文本-图像对齐度在553个提示上保持稳定。
  • 感知变化具有焦点特异性,远超随机遮蔽效果,适合研究生成机制者阅读。

注意力焦点——接收过多注意力质量的标记——在自回归语言模型中被认为功能重要,但在扩散变换器中的作用尚不明确。本文针对文生图扩散模型开展因果分析,动态识别每个时间步的主导注意力接收者,并通过无训练的配对干预在得分和值路径上抑制这些焦点。在553个GenEval提示(基于Stable Diffusion 3,SDXL验证)上,移除这些焦点在k=1时未导致文本-图像对齐(CLIP-T)或偏好代理(ImageReward、HPS-v2)下降;仅在更强干预(k≥10)时,HPS-v2出现依赖指标的边界,而CLIP-T始终稳健。然而,抑制引起的感知变化具有显著的焦点特异性,幅度约为等预算随机遮蔽的6倍,揭示了扩散变换器中轨迹扰动与语义对齐之间的实证分离。

原文摘要 · Abstract (English)

Attention sinks -- tokens that receive disproportionate attention mass -- are assumed to be functionally important in autoregressive language models, but their role in diffusion transformers remains unclear. We present a causal analysis in text-to-image diffusion, dynamically identifying dominant attention recipients per timestep and suppressing them via paired, training-free interventions on the score and value paths. Across 553 GenEval prompts on Stable Diffusion~3 (with SDXL corroboration), removing these sinks does not degrade text-image alignment (CLIP-T) or preference proxies (ImageReward, HPS-v2) at $k{=}1$; only under stronger interventions ($k\!\geq\!10$) does HPS-v2 exhibit a metric-dependent boundary, while CLIP-T remains robust throughout. The perceptual shifts induced by suppression are nonetheless \emph{sink-specific} -- $\sim\!6\times$ larger than equal-budget random masking -- revealing an empirical dissociation between trajectory-level perturbation and \emph{semantic alignment} in diffusion transformers. \footnote{Code available at https://github.com/wfz666/ICML26-attention-sink.}

扩散模型注意力机制生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。