arXiv:2412.07761cs.CV2024-12CVPR被引 25

用预训练扩散模型解决事件相机视频插帧难题

Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation

论文配图:Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
图 1 · 摘自论文原文
  • 将互联网级数据预训练的视频扩散模型迁移到事件相机插帧
  • 在真实数据集上超越现有方法,跨摄像头泛化能力更强
  • 无需大量配对数据,适合事件相机视频增强场景

视频帧插值旨在从低帧率视频中恢复出高帧率的连续视频。然而,由于帧间运动过大,该问题在缺乏额外引导的情况下是病态的。事件相机视频插值(EVFI)通过利用稀疏、高时间分辨率的事件信号作为运动引导,显著优于仅依赖图像帧的方法。然而,现有方法受限于有限的成对事件-帧训练数据,严重制约了性能与泛化能力。本文提出将大规模互联网数据预训练的视频扩散模型迁移到EVFI任务中,克服数据稀缺问题。我们在真实世界数据集上验证了方法的有效性,包括一个新引入的数据集。实验结果表明,该方法在多个数据集上均优于现有方法,并在跨相机场景下展现出更优的泛化能力。

原文摘要 · Abstract (English)

Video Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, the large motion between frames makes this problem ill-posed. Event-based Video Frame Interpolation (EVFI) addresses this challenge by using sparse, high-temporal-resolution event measurements as motion guidance. This guidance allows EVFI methods to significantly outperform frame-only methods. However, to date, EVFI methods have relied on a limited set of paired event-frame training data, severely limiting their performance and generalization capabilities. In this work, we overcome the limited data challenge by adapting pre-trained video diffusion models trained on internet-scale datasets to EVFI. We experimentally validate our approach on real-world EVFI datasets, including a new one that we introduce. Our method outperforms existing methods and generalizes across cameras far better than existing approaches.

视频插帧事件相机扩散模型迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。