arXiv:2506.05301cs.CV2025-06被引 45

一拍即合的视频修复:单步完成高清视频质量提升

SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training

  • 用自适应窗口注意力机制应对高分辨率视频修复
  • 单步推理下性能媲美甚至超越现有方法
  • 适合追求高效高清视频修复的研究与工程人员

基于扩散模型的视频修复(VR)虽显著提升了视觉质量,但推理计算成本过高。尽管已有基于知识蒸馏的一步图像修复方法,但将其拓展至视频领域仍具挑战且研究不足,尤其在真实场景下的高分辨率视频修复中。本文提出一种一步式扩散视频修复模型SeedVR2,通过对抗性后训练直接对真实数据进行优化。为解决单步高分辨率修复中的难题,我们改进了模型架构与训练流程:引入自适应窗口注意力机制,动态调整窗口大小以匹配输出分辨率,避免预设窗口尺寸导致的不一致问题;并验证了多种损失函数的有效性,包括一种不显著影响训练效率的特征匹配损失。大量实验表明,SeedVR2在单步推理下可达到甚至超越现有方法的性能。

原文摘要 · Abstract (English)

Recent advances in diffusion-based video restoration (VR) demonstrate significant improvement in visual quality, yet yield a prohibitive computational cost during inference. While several distillation-based approaches have exhibited the potential of one-step image restoration, extending existing approaches to VR remains challenging and underexplored, particularly when dealing with high-resolution video in real-world settings. In this work, we propose a one-step diffusion-based VR model, termed as SeedVR2, which performs adversarial VR training against real data. To handle the challenging high-resolution VR within a single step, we introduce several enhancements to both model architecture and training procedures. Specifically, an adaptive window attention mechanism is proposed, where the window size is dynamically adjusted to fit the output resolutions, avoiding window inconsistency observed under high-resolution VR using window attention with a predefined window size. To stabilize and improve the adversarial post-training towards VR, we further verify the effectiveness of a series of losses, including a proposed feature matching loss without significantly sacrificing training efficiency. Extensive experiments show that SeedVR2 can achieve comparable or even better performance compared with existing VR approaches in a single step.

视频修复扩散模型单步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。