arXiv:2502.07508cs.CV2025-02被引 10

无需训练即可提升扩散模型生成视频的连贯性与质量。

Enhance-A-Video: Better Generated Video for Free

  • 通过非对角线时间注意力增强帧间关联性。
  • 在多个视频生成模型上显著提升时序一致性与画质。
  • 无需重训练,通用性强,适合快速部署。

基于DiT的视频生成已取得显著成果,但对现有模型的增强研究仍较匮乏。本文提出一种无需训练的视频增强方法 Enhance-A-Video,核心思路是基于非对角线时间注意力分布增强帧间相关性。由于设计简洁,该方法可直接应用于多数DiT-based视频生成框架,无需任何重新训练或微调。在多种DiT-based视频生成模型上,该方法均展现出在时序一致性和视觉质量上的显著提升。我们希望此项工作能激发未来在视频生成增强方向的探索。

原文摘要 · Abstract (English)

DiT-based video generation has achieved remarkable results, but research into enhancing existing models remains relatively unexplored. In this work, we introduce a training-free approach to enhance the coherence and quality of DiT-based generated videos, named Enhance-A-Video. The core idea is enhancing the cross-frame correlations based on non-diagonal temporal attention distributions. Thanks to its simple design, our approach can be easily applied to most DiT-based video generation frameworks without any retraining or fine-tuning. Across various DiT-based video generation models, our approach demonstrates promising improvements in both temporal consistency and visual quality. We hope this research can inspire future explorations in video generation enhancement.

视频生成扩散模型时序一致性零样本增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。