arXiv:2503.14863cs.CV2025-03AAAI被引 4

用预训练扩散模型修复视频,提升画质与时间一致性。

Temporal-Consistent Video Restoration with Pre-trained Diffusion Models

  • 将反向扩散过程视为函数,直接在种子空间参数化视频帧
  • 在多个虚拟现实任务中实现更优的视觉质量和时间一致性
  • 适合需要高质量视频修复的应用场景

视频修复(VR)旨在从退化视频中恢复高质量内容。尽管基于预训练扩散模型(DMs)的零样本修复方法展现出良好前景,但仍存在反向扩散中的近似误差和时间一致性不足的问题。此外,处理3D视频数据使修复过程计算成本高昂。本文将扩散模型的反向过程视为函数,提出一种新的最大后验(MAP)框架,直接在模型的种子空间中参数化视频帧,从而消除近似误差。同时引入促进双层次时间一致性的策略:利用种子空间中的聚类结构实现语义一致性,通过带光流优化的渐进形变实现像素级一致性。在多个虚拟现实任务上的大量实验表明,该方法在视觉质量与时间一致性方面均优于现有最先进水平。

原文摘要 · Abstract (English)

Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they suffer from approximation errors during reverse diffusion and insufficient temporal consistency. Moreover, dealing with 3D video data, VR is inherently computationally intensive. In this paper, we advocate viewing the reverse process in DMs as a function and present a novel Maximum a Posterior (MAP) framework that directly parameterizes video frames in the seed space of DMs, eliminating approximation errors. We also introduce strategies to promote bilevel temporal consistency: semantic consistency by leveraging clustering structures in the seed space, and pixel-level consistency by progressive warping with optical flow refinements. Extensive experiments on multiple virtual reality tasks demonstrate superior visual quality and temporal consistency achieved by our method compared to the state-of-the-art.

视频修复扩散模型时间一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。