arXiv:2501.12267cs.CV2025-01中稿 · WACV 2025

无需训练即可生成连贯多样的视频修复结果

VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models

  • 利用光流引导从参考帧提取有效像素,约束噪声优化
  • 在大区域缺失时仍保持时空一致性,显著减少伪影
  • 支持多样化生成,适合需要高自由度修复的场景

近期视频修复方法通过光流引导像素传播,在图像或特征空间中取得进展。然而,当掩码区域过大、中心缺乏对应像素时,会产生严重伪影。扩散模型虽在图像修复中表现优异,但难以直接用于视频以保证时间连贯性。本文提出无需训练的VipDiff框架,通过反向扩散过程条件化扩散模型,在不依赖训练数据或微调预训练模型的前提下,实现时空一致的视频修复。VipDiff利用光流提取参考帧中的有效像素作为约束,优化随机采样的高斯噪声,并结合生成结果进行后续像素传播与条件生成。同时支持不同噪声样本下的多样化输出。实验表明,VipDiff在时空一致性与保真度上均显著优于现有先进方法。

原文摘要 · Abstract (English)

Recent video inpainting methods have achieved encouraging improvements by leveraging optical flow to guide pixel propagation from reference frames either in the image space or feature space. However, they would produce severe artifacts in the mask center when the masked area is too large and no pixel correspondences can be found for the center. Recently, diffusion models have demonstrated impressive performance in generating diverse and high-quality images, and have been exploited in a number of works for image inpainting. These methods, however, cannot be applied directly to videos to produce temporal-coherent inpainting results. In this paper, we propose a training-free framework, named VipDiff, for conditioning diffusion model on the reverse diffusion process to produce temporal-coherent inpainting results without requiring any training data or fine-tuning the pre-trained diffusion models. VipDiff takes optical flow as guidance to extract valid pixels from reference frames to serve as constraints in optimizing the randomly sampled Gaussian noise, and uses the generated results for further pixel propagation and conditional generation. VipDiff also allows for generating diverse video inpainting results over different sampled noise. Experiments demonstrate that VipDiff can largely outperform state-of-the-art video inpainting methods in terms of both spatial-temporal coherence and fidelity.

视频修复扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。