arXiv:2501.01320cs.CV2025-01CVPR被引 70

提出可处理任意长度与分辨率的视频修复模型,解决真实场景下时序一致性难题。

SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration

  • 采用偏移窗口注意力机制,高效处理长视频序列。
  • 支持空间与时间边界处可变窗口,突破传统注意力分辨率限制。
  • 适配真实世界、合成数据及AI生成视频,通用性强。

视频修复在保持保真度的同时,需从真实世界未知退化中恢复时序一致的细节,面临巨大挑战。尽管基于扩散模型的方法取得进展,但生成能力与采样效率仍受限。本文提出SeedVR,一种用于任意长度和分辨率视频修复的扩散变压器。其核心是偏移窗口注意力机制,有效处理长视频序列;并在空间与时间边界支持可变尺寸窗口,克服传统窗口注意力的分辨率瓶颈。结合因果视频自编码器、图像与视频混合训练及渐进式训练等现代技术,SeedVR在合成数据、真实世界数据及AI生成视频上均达到领先性能,实验表明其在通用视频修复任务上显著优于现有方法。

原文摘要 · Abstract (English)

Video restoration poses non-trivial challenges in maintaining fidelity while recovering temporally consistent details from unknown degradations in the wild. Despite recent advances in diffusion-based restoration, these methods often face limitations in generation capability and sampling efficiency. In this work, we present SeedVR, a diffusion transformer designed to handle real-world video restoration with arbitrary length and resolution. The core design of SeedVR lies in the shifted window attention that facilitates effective restoration on long video sequences. SeedVR further supports variable-sized windows near the boundary of both spatial and temporal dimensions, overcoming the resolution constraints of traditional window attention. Equipped with contemporary practices, including causal video autoencoder, mixed image and video training, and progressive training, SeedVR achieves highly-competitive performance on both synthetic and real-world benchmarks, as well as AI-generated videos. Extensive experiments demonstrate SeedVR's superiority over existing methods for generic video restoration.

视频修复扩散模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。