arXiv:2510.11538cs.CV2025-10被引 7

发现扩散Transformer中大量激活是细节生成关键,提出无需训练的增强方法

Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers

  • 通过破坏大规模激活构造退化模型,引导原模型提升细节质量
  • 在SD3、SD3.5、Flux等模型上均显著改善细粒度细节表现
  • 可无缝融合分类器自由指导,适合追求高细节图像生成的研究者

扩散Transformer(DiTs)作为视觉生成的强大骨干网络,其内部特征图中存在大量激活(MAs),但其作用尚不明确。本研究系统分析发现,这些激活分布于所有空间标记,并受输入时间步嵌入调控。重要的是,大规模激活对输出整体语义影响小,却在局部细节合成中起关键作用。基于此,我们提出无需训练的细节引导(DG)策略,通过破坏MAs构建一个细节缺失的退化模型,用以引导原模型优化细节生成。DG可与分类器自由指导(CFG)结合,进一步细化细粒度细节。大量实验表明,该方法在多个预训练DiTs(如SD3、SD3.5和Flux)上持续提升细节质量。

原文摘要 · Abstract (English)

Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feature maps, yet their function remains poorly understood. In this work, we systematically investigate these activations to elucidate their role in visual generation. We found that these massive activations occur across all spatial tokens, and their distribution is modulated by the input timestep embeddings. Importantly, our investigations further demonstrate that these massive activations play a key role in local detail synthesis, while having minimal impact on the overall semantic content of output. Building on these insights, we propose \textbf{D}etail \textbf{G}uidance (\textbf{DG}), a MAs-driven, training-free self-guidance strategy to explicitly enhance local detail fidelity for DiTs. Specifically, DG constructs a degraded ``detail-deficient'' model by disrupting MAs and leverages it to guide the original network toward higher-quality detail synthesis. Our DG can seamlessly integrate with Classifier-Free Guidance (CFG), enabling further refinements of fine-grained details. Extensive experiments demonstrate that our DG consistently improves fine-grained detail quality across various pre-trained DiTs (\eg, SD3, SD3.5, and Flux).

扩散模型细节生成自指导Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。