通过调控大规模激活提升扩散Transformer的生成与理解能力
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

- 发现扩散Transformer中激活集中在特定特征维度,由去噪时间步主导
- 提出EMA框架,利用激活信号同时增强图像细节生成和特征区分度
- 无需训练即可改进生成质量与视觉表征,适用于图像生成与理解任务
在基于Transformer的扩散模型(DiTs)中,大规模激活(MAs)普遍存在,但其结构与功能仍不明确。本文系统分析代表性DiTs中的MAs,发现其在图像令牌空间上分布,却集中于少数固定特征维度。这些维度与AdaLN残差缩放因子高度对齐,主要受去噪时间步而非文本条件调控。该结构导致双重任务效应:生成时MAs对精细细节合成至关重要,但对全局语义影响有限;理解时其高幅方向使原始特征在空间令牌间过于相似,削弱了密集特征区分能力。基于此,我们提出免训练的EMA框架,以MAs为统一调制信号,提升DiTs的生成与表征能力。生成方面,引入MA驱动的细节引导(DG),通过抑制MA维度生成缺乏细节的反事实预测,引导采样聚焦细粒度视觉特征,并支持部分前向推理、无分类器引导集成及局部细节优化;理解方面,提出MA调制的表示提取(MREP),利用预训练的AdaLN通道调制降低MA方向主导性,并拼接空间归一化后的MA图以保留有效空间结构。大量实验表明,EMA能持续提升DiTs的生成质量与表征能力。
原文摘要 · Abstract (English)
Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiTs) remain insufficiently understood. In this work, we systematically analyze MAs in representative DiTs and find that they are spatially distributed across image tokens while concentrated in a small set of fixed feature dimensions. We further show that these dimensions are closely aligned with AdaLN residual scaling factors and are primarily modulated by the denoising timestep rather than text conditions. This structure leads to two task-dependent effects: for generation, MAs are critical for fine-grained detail synthesis while having limited influence on global semantics; for understanding, their shared high-magnitude directions make raw DiT features overly similar across spatial tokens and weaken dense feature discrimination. Based on these findings, we introduce Eliciting Massive Activation (EMA), a training-free framework that leverages Massive Activations (MAs) as a unified modulation signal to improve both generative and representational capabilities of DiTs. For generation, EMA proposes MA-driven Detail G}uidance (DG), which suppresses MA dimensions to construct a detail-deficient counterfactual prediction and guides sampling toward finer visual details. DG further supports efficient partial-forward inference, integration with classifier-free guidance, and token-level Local DG for refining selected image regions. For understanding, EMA introduces MA-modulated REPresentation extraction (MREP), which uses pretrained AdaLN channel-wise modulation to reduce MA directional dominance and concatenates spatially normalized MA maps to preserve useful spatial structure. Extensive experiments demonstrate that EMA consistently improves both the generation quality and representation capability of DiTs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。