SEGA动态调节注意力,让扩散模型生成更高清图像更清晰。
SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers

- 根据潜在空间频率结构,动态调整旋转位置编码的注意力
- 在多个高分辨率下均提升图像结构与细节质量
- 适合需要高清生成但无法重新训练的场景
扩散变压器(DiTs)已成为文本到图像生成的主流架构,但在超出训练分辨率范围时性能下降。现有无需训练的方法通过修改推理时的注意力行为来缓解,常结合旋转位置编码(RoPE)外推与注意力缩放。然而这些方法对不同频率特性的RoPE组件采用统一且内容无关的缩放,导致全局结构保持与细粒度恢复之间的权衡。我们提出SEGA,一种无需训练的方法,根据每个去噪步骤中潜在表示的空间频率结构,动态调整各RoPE组件的注意力缩放。该自适应缩放同时提升结构连贯性与细节保真度。实验表明,SEGA在多个目标分辨率上持续优于当前最优的无需训练基线。
原文摘要 · Abstract (English)
Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resolutions beyond their training range. Existing training-free approaches mitigate this by modifying inference-time attention behavior, often through Rotary Position Embeddings (RoPE) extrapolation combined with attention scaling. However, these strategies apply a uniform and content-agnostic scaling across RoPE components with distinct frequency characteristics, inducing a trade-off between preserving global structure and recovering fine detail. We introduce SEGA, a training-free method that dynamically scales attention across RoPE components according to the latent's spatial-frequency structure at each denoising step. This adaptive scaling improves both structural coherence and fine-detail fidelity. Experiments show that SEGA consistently improves high-resolution synthesis across multiple target resolutions, outperforming state-of-the-art training-free baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。