用视频扩散模型解决图像修复时的帧间闪烁问题
Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models
- 融合同源与异源视频扩散模型的潜在表示
- 在多个数据集上显著减少帧间闪烁,提升时序一致性
- 无需训练,可通用适配各类图像修复方法
尽管基于扩散模型的零样本图像修复与增强方法已取得显著进展,但将其应用于视频修复或增强时,会引发严重的时序闪烁问题。本文提出首个利用快速发展的视频扩散模型辅助图像方法以维持更高时序一致性的框架。通过引入同源潜在融合、异源潜在融合以及基于思维链(COT)的融合比例策略,有效结合同源与异源文本到视频扩散模型,弥补图像方法的不足。此外,提出时序强化后处理模块,进一步利用图像到视频扩散模型提升时序一致性。所提方法无需训练,可适配任意基于扩散的图像修复与增强方法。实验结果表明其在多个数据集上均优于现有方法。
原文摘要 · Abstract (English)
Although diffusion-based zero-shot image restoration and enhancement methods have achieved great success, applying them to video restoration or enhancement will lead to severe temporal flickering. In this paper, we propose the first framework that utilizes the rapidly-developed video diffusion model to assist the image-based method in maintaining more temporal consistency for zero-shot video restoration and enhancement. We propose homologous latents fusion, heterogenous latents fusion, and a COT-based fusion ratio strategy to utilize both homologous and heterogenous text-to-video diffusion models to complement the image method. Moreover, we propose temporal-strengthening post-processing to utilize the image-to-video diffusion model to further improve temporal consistency. Our method is training-free and can be applied to any diffusion-based image restoration and enhancement methods. Experimental results demonstrate the superiority of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。