arXiv:2603.02531cs.LGcs.AI2026-03

提出无需训练的注意力引导方法,提升扩散模型生成质量。

Geometry-Aware Attention Guidance for Diffusion Models via Modern Hopfield Dynamics

  • 基于现代霍普菲尔德动力学分析注意力差异,设计方向一致的加速信号。
  • 在少步采样下显著提升生成质量,支持多种模型架构和采样方式。
  • 无需额外训练,可直接部署于UNet、MMDiT等主流模型中。

无分类器引导(CFG)虽能提升扩散模型的样本质量,但其双阶段推理与对零条件训练的依赖限制了其在少步采样场景下的应用。注意力空间引导作为互补范式应运而生,但现有稀疏-稠密注意力引导为何有效仍不明确。本文通过现代霍普菲尔德动力学分析注意力外推,证明在共享条件下的稀疏-稠密差异具有两个方向性特性,共同构成方向一致的加速信号。基于此,我们提出几何感知注意力引导(GAG),一种无需训练、即插即用的外推规则:将差异分解为沿检索方向的平行与正交分量,增强对齐收敛方向的成分,抑制离流形噪声;稳定性源于弱压缩性质。进一步将该外推解释为注意力空间的一阶安德森加速,统一了现有注意力外推方法视角。GAG为通用方法,跨架构(UNet、MMDiT)与采样场景(多步、少步)均表现良好,在包括FLUX.1、FLUX.2和Qwen-Image在内的多种骨干模型上持续提升生成质量,计算开销极低。

原文摘要 · Abstract (English)

Classifier-Free Guidance (CFG) improves sample quality in diffusion models, but its dual-pass inference and reliance on null-condition training limit its use in few-step regimes. Attention-space guidance has emerged as a complementary paradigm that addresses this gap, yet why prior sparse-vs-dense attention guidance works remains elusive. We address this by analyzing attention extrapolation through Modern Hopfield dynamics, proving two directional properties of the sparse-dense discrepancy under shared conditioning that together certify it as a directionally consistent acceleration signal. Building on this, we propose Geometry-Aware Attention Guidance (GAG), a training-free, plug-and-play extrapolation rule that decomposes the discrepancy into parallel and orthogonal components relative to the retrieval direction, amplifying the convergence-aligned component while suppressing off-manifold noise; stability follows from a weak contraction property. We further provide an interpretation of this extrapolation as first-order Anderson Acceleration in attention space, offering a unified perspective on attention extrapolation methods. GAG is a universal method that generalizes across architectures (UNet, MMDiT) and sampling regimes (multi-step, few-step), consistently improving generation quality on diverse backbones, including FLUX.1, the recent FLUX.2, and Qwen-Image, with minimal computational overhead.

扩散模型注意力机制生成质量少步采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。