发现视频扩散模型中稀有激活峰值可调控生成质量,无需训练即可提升视频清晰度与连贯性。
Steering Video Diffusion Transformers with Massive Activations
- 识别出视频扩散模型中特定位置的高幅度稀有激活,其分布具有结构规律。
- 通过调节这些激活强度,可显著提升文本生成视频的质量与时间一致性。
- 提出无训练干预的激活引导方法,适合快速优化现有视频生成模型。
本文研究视频扩散变换器(DiTs)中罕见的高幅值激活(Massive Activations, MAs),它们集中在少数固定隐藏维度上。我们发现MAs在首帧标记和潜变量帧的空间边界处达到峰值,且这一模式在早期去噪阶段最为明显。该现象源于视频VAE的编码不对称性:因果时间填充与零空间填充导致首帧及帧边界内容负载较低,而高MAs恰好对应这些低负载区域。分析中间表示发现,MAs充当残差计算的隐式缩放器:增强MAs会抑制自注意力与前馈更新,移除则放大更新。这表明视频DiTs以不均匀方式部署残差缩放,最强阻尼位于编码不对称的结构位置。基于此,我们提出结构化激活引导(STAS),一种无需额外前向传播的训练自由技术,在早期去噪阶段将观察到的结构性位置上的MAs导向模型生成的参考值。STAS能一致提升多款文本到视频模型的视频质量与时间连贯性,开销极小。
原文摘要 · Abstract (English)
In this work, we study the role of Massive Activations (MAs), which are rare, high-magnitude spikes confined to a few fixed hidden dimensions in video diffusion transformers (DiTs). We uncover a structured positional hierarchy: MA magnitudes peak at first-frame tokens and recur at the spatial boundary tokens of latent frames, with this pattern being most pronounced during early denoising. We trace this organization to an encoding asymmetry of the video VAEs, whose causal temporal padding and zero spatial padding cause the first latent frame and frame borders to carry reduced content load. Elevated MAs consistently align with these lower-content structural positions. To understand their function, we analyze intermediate representations and find that MAs act as implicit rescalers of residual computation: enlarging MAs suppresses the corresponding self-attention and feed-forward updates, while erasing them amplifies these updates. Together, these observations suggest that MAs serves as a token-level rescaler of residual computation, which video DiTs deploy unevenly, placing the strongest damping at the encoding-asymmetric structural positions. Motivated by this native rescaling behavior, we propose Structured Activation Steering (STAS), a training-free technique that steers MAs at the observed structural positions toward a scaled, model-derived reference during early denoising. STAS requires no additional forward passes and consistently improves video quality and temporal coherence across text-to-video models with negligible overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。