用改进的视频扩散模型实现更真实、细节更丰富的去模糊
DIVD: Deblurring with Improved Video Diffusion Model
- 基于视频扩散模型,利用帧间相关性提升去模糊效果
- 在感知指标上超越现有方法,细节保留更出色
- 首个将扩散模型用于视频去模糊的工作,适合图像质量敏感场景
视频去模糊因模糊成因复杂(如相机抖动与物体运动)而极具挑战。以往方法多依赖PSNR等失真度量,但与人类感知关联弱,重建结果缺乏真实感。扩散模型在图像与视频生成领域表现优异,尤其在真实性与感知效果上领先。然而,由于计算复杂性和适配难题,视频扩散模型在去模糊任务中的潜力尚不明确。为此,本文提出专用于视频去模糊的扩散模型,通过挖掘相邻帧间的高度相关性并解决时间错位问题,引入多项改进。实验表明,该模型在多种感知指标上达到最优,同时保持了良好的失真指标(如PSNR),显著提升了图像细节还原能力。据我们所知,这是首次将扩散模型成功应用于视频去模糊任务,突破了传统方法的局限。
原文摘要 · Abstract (English)
Video deblurring presents a considerable challenge owing to the complexity of blur, which frequently results from a combination of camera shakes, and object motions. In the field of video deblurring, many previous works have primarily concentrated on distortion-based metrics, such as PSNR. However, this approach often results in a weak correlation with human perception and yields reconstructions that lack realism. Diffusion models and video diffusion models have respectively excelled in the fields of image and video generation, particularly achieving remarkable results in terms of image authenticity and realistic perception. However, due to the computational complexity and challenges inherent in adapting diffusion models, there is still uncertainty regarding the potential of video diffusion models in video deblurring tasks. To explore the viability of video diffusion models in the task of video deblurring, we introduce a diffusion model specifically for this purpose. In this field, leveraging highly correlated information between adjacent frames and addressing the challenge of temporal misalignment are crucial research directions. To tackle these challenges, many improvements based on the video diffusion model are introduced in this work. As a result, our model outperforms existing models and achieves state-of-the-art results on a range of perceptual metrics. Our model preserves a significant amount of detail in the images while maintaining competitive distortion metrics. Furthermore, to the best of our knowledge, this is the first time the diffusion model has been applied in video deblurring to overcome the limitations mentioned above.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。