arXiv:2501.02269cs.CV2025-01被引 5

首个统一视频修复扩散模型,一模型搞定多种画质问题。

TDM: Temporally-Consistent Diffusion Model for All-in-One Real-World Video Restoration

  • 用预训练扩散模型+微调ControlNet,统一处理多种视频退化。
  • 结合DDIM反演与滑动窗口跨帧注意力,提升时序一致性。
  • 支持多任务扩展,适合实际场景中复杂退化视频修复。

本文提出首个基于扩散模型的全功能视频修复方法,利用预训练的Stable Diffusion与微调的ControlNet,仅用一个统一模型即可恢复多种类型的视频退化问题,突破了传统方法需为每类任务单独建模的局限。核心贡献包括:采用任务提示引导(TPG)的高效训练策略,支持多样化修复任务;结合去噪扩散隐式模型(DDIM)反演与新型滑动窗口跨帧注意力(SW-CFA)机制的推理策略,显著增强内容保真度与时序一致性;以及可扩展的全流程架构,实现真正意义上的“一站式”修复。在五个视频修复任务上的大量实验表明,该方法在真实世界视频上具有更强泛化能力与更优的时序一致性表现,优于现有最先进方法。本工作推动了视频修复向统一、通用解决方案迈进。

原文摘要 · Abstract (English)

In this paper, we propose the first diffusion-based all-in-one video restoration method that utilizes the power of a pre-trained Stable Diffusion and a fine-tuned ControlNet. Our method can restore various types of video degradation with a single unified model, overcoming the limitation of standard methods that require specific models for each restoration task. Our contributions include an efficient training strategy with Task Prompt Guidance (TPG) for diverse restoration tasks, an inference strategy that combines Denoising Diffusion Implicit Models~(DDIM) inversion with a novel Sliding Window Cross-Frame Attention (SW-CFA) mechanism for enhanced content preservation and temporal consistency, and a scalable pipeline that makes our method all-in-one to adapt to different video restoration tasks. Through extensive experiments on five video restoration tasks, we demonstrate the superiority of our method in generalization capability to real-world videos and temporal consistency preservation over existing state-of-the-art methods. Our method advances the video restoration task by providing a unified solution that enhances video quality across multiple applications.

视频修复扩散模型时序一致统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。