arXiv:2506.10978cs.CVcs.AI2025-06NeurIPS被引 5

提出细粒度注意力扰动方法,实现对图像风格与质量的精准控制。

Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models

  • 按注意力头级别设计扰动,定位影响结构、风格、纹理的关键模块。
  • 在Stable Diffusion 3和FLUX.1上提升生成质量,减少模糊与伪影。
  • 支持用户自定义风格组合,适合可控图像生成研究者使用。

扩散模型中的引导方法通过扰动模型构建隐式弱模型,引导生成过程远离该模型。现有注意力扰动方法在无分类器引导场景下表现良好,但缺乏对扰动位置的系统性设计,尤其在注意力计算分布于多层的扩散变压器(DiT)架构中。本文研究了从层到具体注意力头的扰动粒度,发现特定注意力头分别控制结构、风格和纹理质量等视觉概念。基于此,提出HeadHunter框架,可迭代选择符合用户目标的注意力头,实现对生成质量与视觉属性的细粒度控制。同时引入SoftPAG,通过线性插值将选中头的注意力图向单位矩阵过渡,提供连续可调的扰动强度,抑制伪影。该方法缓解了传统层级扰动的过度平滑问题,并可通过组合头部实现特定风格操控。在Stable Diffusion 3与FLUX.1等大规模DiT模型上验证,显著提升通用质量与风格引导效果。本工作首次实现扩散模型中注意力扰动的头级别分析,揭示了注意力层内的可解释专业化特性,为有效扰动策略设计提供实践基础。

原文摘要 · Abstract (English)

Recent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-free guidance is not applicable. However, existing attention perturbation methods lack principled approaches for determining where perturbations should be applied, particularly in Diffusion Transformer (DiT) architectures where quality-relevant computations are distributed across layers. In this paper, we investigate the granularity of attention perturbations, ranging from the layer level down to individual attention heads, and discover that specific heads govern distinct visual concepts such as structure, style, and texture quality. Building on this insight, we propose "HeadHunter", a systematic framework for iteratively selecting attention heads that align with user-centric objectives, enabling fine-grained control over generation quality and visual attributes. In addition, we introduce SoftPAG, which linearly interpolates each selected head's attention map toward an identity matrix, providing a continuous knob to tune perturbation strength and suppress artifacts. Our approach not only mitigates the oversmoothing issues of existing layer-level perturbation but also enables targeted manipulation of specific visual styles through compositional head selection. We validate our method on modern large-scale DiT-based text-to-image models including Stable Diffusion 3 and FLUX.1, demonstrating superior performance in both general quality enhancement and style-specific guidance. Our work provides the first head-level analysis of attention perturbation in diffusion models, uncovering interpretable specialization within attention layers and enabling practical design of effective perturbation strategies.

扩散模型注意力机制图像生成可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。