arXiv:2504.12027cs.CV2025-04被引 4

解析视频生成中注意力机制如何影响画质与连贯性

Understanding Attention Mechanism in Video Diffusion Models

  • 通过信息论方法分析时空注意力块的作用机制
  • 高熵注意力图与视频质量正相关,低熵则关联帧内结构
  • 提出轻量级注意力调控方法,支持文本引导编辑

文生视频(T2V)模型如OpenAI的Sora因其从文本提示生成高质量视频的能力而备受关注。在基于扩散的T2V模型中,注意力机制是关键组件。然而,其学习到的中间特征尚不明确,且注意力模块如何影响视频合成中的图像质量与时间一致性仍不清楚。本文采用信息论方法对T2V模型的时空注意力块进行深入扰动分析。结果表明,时空注意力图不仅影响视频的时间布局,还决定时空元素的复杂度与美学质量。值得注意的是,高熵注意力图常与更优视频质量相关,而低熵图则与帧内结构有关。基于此,我们提出两种新方法以提升视频质量并实现文本引导编辑,仅通过轻量级操作注意力矩阵即可完成。在多个数据集上的实验验证了方法的有效性。

原文摘要 · Abstract (English)

Text-to-video (T2V) synthesis models, such as OpenAI's Sora, have garnered significant attention due to their ability to generate high-quality videos from a text prompt. In diffusion-based T2V models, the attention mechanism is a critical component. However, it remains unclear what intermediate features are learned and how attention blocks in T2V models affect various aspects of video synthesis, such as image quality and temporal consistency. In this paper, we conduct an in-depth perturbation analysis of the spatial and temporal attention blocks of T2V models using an information-theoretic approach. Our results indicate that temporal and spatial attention maps affect not only the timing and layout of the videos but also the complexity of spatiotemporal elements and the aesthetic quality of the synthesized videos. Notably, high-entropy attention maps are often key elements linked to superior video quality, whereas low-entropy attention maps are associated with the video's intra-frame structure. Based on our findings, we propose two novel methods to enhance video quality and enable text-guided video editing. These methods rely entirely on lightweight manipulation of the attention matrices in T2V models. The efficacy and effectiveness of our methods are further validated through experimental evaluation across multiple datasets.

视频生成注意力机制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。