用文本生成模型实现零样本视频修复,解决闪烁问题并提速近三分之二
Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

- 结合文本与多模态参考,用扩散模型实现无需训练的视频修复
- 通过双提示调优和采样,推理速度降至原来的1/3,且时间一致性更强
- 适合需要快速修复视频且无标注数据的场景,如老片修复、监控视频增强
基于文本到图像潜在扩散模型的零样本图像修复方法在无需训练的情况下已取得显著成果。然而将其应用于视频修复时会出现严重的帧间闪烁问题。本文提出一种新型零样本视频修复与增强框架,结合文本到图像潜在扩散模型与多模态参考。通过提出的双提示调优反演与采样策略,推理时间可降至原方法的约1/3,性能与时间一致性均显著提升。通过纹理感知的视频标记合并机制,进一步利用帧间时序相关性以改善时间一致性。此外,提出参考自注意力与参考标记合并机制,支持图像参考输入。实验表明,所提方法在恢复高质量且时序一致的视频方面具有明显优势。
原文摘要 · Abstract (English)
Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without training. However, applying them to video restoration will result in severe temporal flickering. In this paper, we propose a novel framework for zero-shot video restoration and enhancement which uses a text-to-image latent diffusion model and multi-modal references. Through the proposed dual prompt tuning inversion and sampling, the inference time can be reduced to nearly 1/3 of the original. The performance and temporal consistency can be also significantly stregthened. By using the proposed texture-aware video token merging, the temporal correlation between frames can be further utilized to improve the temporal consistency. We futher propose the referenced self-attention and referenced token merging to support image reference. Experimental results demonstrate the superiority of the proposed method in restoring and enhancing temporally consistent videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。