arXiv:2511.07499cs.CVcs.AI2025-11AAAI被引 2

用对抗性注意力引导提升扩散模型采样可靠性,无需重训练。

Toward the Frontiers of Reliable Diffusion Sampling via Adversarial Sinkhorn Attention Guidance

  • 基于最优传输思想,在自注意力中注入对抗性代价,破坏错误对齐。
  • 在文生图任务中显著提升条件与无条件样本质量,改善可控性与保真度。
  • 轻量级插件式设计,适用于IP-Adapter、ControlNet等下游应用。

扩散模型在使用分类器自由引导(CFG)等引导方法时表现出强大的生成能力,这些方法通过修改采样轨迹提升输出质量。然而,这类方法通常通过启发式扰动函数(如恒等混合或模糊条件)故意劣化另一输出(如无条件输出),缺乏理论基础且依赖人工设计的失真方式。本文提出对抗性Sinkhorn注意力引导(ASAG),将扩散模型中的注意力得分重新解释为最优传输问题,并通过Sinkhorn算法有意干扰传输成本。不同于简单破坏注意力机制,ASAG在自注意力层中引入对抗性代价,降低查询与键之间的像素级相似性,从而削弱误导性注意力对齐,提升条件与无条件样本质量。ASAG在文生图任务中表现一致提升,增强下游应用(如IP-Adapter和ControlNet)的可控性与保真度。该方法轻量、可即插即用,无需模型重训练。

原文摘要 · Abstract (English)

Diffusion models have demonstrated strong generative performance when using guidance methods such as classifier-free guidance (CFG), which enhance output quality by modifying the sampling trajectory. These methods typically improve a target output by intentionally degrading another, often the unconditional output, using heuristic perturbation functions such as identity mixing or blurred conditions. However, these approaches lack a principled foundation and rely on manually designed distortions. In this work, we propose Adversarial Sinkhorn Attention Guidance (ASAG), a novel method that reinterprets attention scores in diffusion models through the lens of optimal transport and intentionally disrupt the transport cost via Sinkhorn algorithm. Instead of naively corrupting the attention mechanism, ASAG injects an adversarial cost within self-attention layers to reduce pixel-wise similarity between queries and keys. This deliberate degradation weakens misleading attention alignments and leads to improved conditional and unconditional sample quality. ASAG shows consistent improvements in text-to-image diffusion, and enhances controllability and fidelity in downstream applications such as IP-Adapter and ControlNet. The method is lightweight, plug-and-play, and improves reliability without requiring any model retraining.

扩散模型注意力机制生成质量轻量部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。