arXiv:2504.10317cs.CV2025-04被引 14

揭示视频扩散模型中注意力机制的三大特性,助力高效视频生成

Analysis of Attention in Video Diffusion Transformers

  • 发现视频注意力具有结构相似性,可迁移自注意力图实现视频编辑
  • 部分层虽看似稀疏但无法压缩,现有稀疏化方法不通用
  • 首次研究注意力汇点,为提升模型效率与质量平衡提供新方向

我们对视频扩散变换器(VDiTs)中的注意力机制进行了深入分析,发现其具备三个关键特性:结构、稀疏性和汇点。结构方面,不同提示下的注意力模式具有相似结构,利用该相似性可通过自注意力图迁移实现视频编辑。稀疏性方面,研究发现现有稀疏化方法并不适用于所有VDiTs,部分看似稀疏的层实际上无法被有效压缩。汇点方面,我们首次系统研究了VDiTs中的注意力汇点,并与语言模型中的汇点进行对比。基于这些发现,我们提出若干未来研究方向,旨在利用这些洞见优化VDiTs在效率与质量之间的权衡表现。

原文摘要 · Abstract (English)

We conduct an in-depth analysis of attention in video diffusion transformers (VDiTs) and report a number of novel findings. We identify three key properties of attention in VDiTs: Structure, Sparsity, and Sinks. Structure: We observe that attention patterns across different VDiTs exhibit similar structure across different prompts, and that we can make use of the similarity of attention patterns to unlock video editing via self-attention map transfer. Sparse: We study attention sparsity in VDiTs, finding that proposed sparsity methods do not work for all VDiTs, because some layers that are seemingly sparse cannot be sparsified. Sinks: We make the first study of attention sinks in VDiTs, comparing and contrasting them to attention sinks in language models. We propose a number of future directions that can make use of our insights to improve the efficiency-quality Pareto frontier for VDiTs.

视频生成注意力机制扩散模型模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。