arXiv:2412.17254cs.CVcs.AI2024-12被引 1

不微调模型就能提升长视频生成一致性,靠频率分析重加权注意力。

Enhancing Long Video Generation Consistency without Tuning

  • 基于短时傅里叶变换重加权时间注意力,提升帧间一致性。
  • 在多提示生成中,提示对齐问题影响插值质量,改进后显著提升。
  • 无需微调,适用于单/多提示场景,适合视频生成研究者使用。

尽管长视频生成已取得显著进展,但生成视频在流畅性与场景过渡方面仍存在明显不一致问题。本文提出时间-频率注意力重加权算法(TiARA),通过离散短时傅里叶变换对注意力分数矩阵进行智能编辑,基于频域分析确保帧间一致性提升,是首个应用于视频扩散模型的频率基方法。针对多提示生成,进一步发现提示对齐程度直接影响提示插值质量,据此提出PromptBlend提示插值流程,系统化对齐提示内容。大量实验证明,该方法在多个基准上均实现稳定且显著的性能提升,无需模型微调即可有效增强生成视频的一致性与连贯性。

原文摘要 · Abstract (English)

Despite the considerable progress achieved in the long video generation problem, there is still significant room to improve the consistency of the generated videos, particularly in terms of their smoothness and transitions between scenes. We address these issues to enhance the consistency and coherence of videos generated with either single or multiple prompts. We propose the Time-frequency based temporal Attention Reweighting Algorithm (TiARA), which judiciously edits the attention score matrix based on the Discrete Short-Time Fourier Transform. This method is supported by a frequency-based analysis, ensuring that the edited attention score matrix achieves improved consistency across frames. It represents the first-of-its-kind for frequency-based methods in video diffusion models. For videos generated by multiple prompts, we further uncover key factors such as the alignment of the prompts affecting prompt interpolation quality. Inspired by our analyses, we propose PromptBlend, an advanced prompt interpolation pipeline that systematically aligns the prompts. Extensive experimental results validate the efficacy of our proposed method, demonstrating consistent and substantial improvements over multiple baselines.

视频生成扩散模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。