arXiv:2503.20268cs.CV2025-03ICCV被引 4

用事件相机引导扩散模型,实现大运动场景的逼真视频插帧

EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolation

  • 融合RGB与事件信号生成运动条件,指导扩散过程
  • 在Prophesee和BSRGB数据集上LPIPS分别提升27.4%和24.1%
  • 适合需要高精度动态建模的视频增强与自动驾驶场景

大运动场景下的视频帧插值仍面临运动模糊问题。尽管事件相机可捕捉高时间分辨率运动信息,现有基于事件的插值方法受限于训练数据少和运动模式复杂。本文提出事件引导视频扩散模型(EGVD),结合预训练稳定视频扩散模型的强大先验与事件相机的精确时序信息。设计多模态运动条件生成器(MMCG),有效融合RGB帧与事件信号以引导扩散过程,生成物理真实的中间帧。采用选择性微调策略,在保留空间建模能力的同时高效融入事件时序信息。引入受最新扩散模型启发的输入输出归一化技术,提升不同噪声水平下的训练稳定性。为增强泛化能力,构建涵盖真实与模拟事件数据的综合数据集。大量实验表明,EGVD在大运动和挑战性光照条件下显著优于现有方法,在Prophesee和BSRGB数据集上感知质量指标LPIPS分别提升27.4%和24.1%,同时保持竞争力保真度。代码与数据集见:https://github.com/OpenImagingLab/EGVD。

原文摘要 · Abstract (English)

Video frame interpolation (VFI) in scenarios with large motion remains challenging due to motion ambiguity between frames. While event cameras can capture high temporal resolution motion information, existing event-based VFI methods struggle with limited training data and complex motion patterns. In this paper, we introduce Event-Guided Video Diffusion Model (EGVD), a novel framework that leverages the powerful priors of pre-trained stable video diffusion models alongside the precise temporal information from event cameras. Our approach features a Multi-modal Motion Condition Generator (MMCG) that effectively integrates RGB frames and event signals to guide the diffusion process, producing physically realistic intermediate frames. We employ a selective fine-tuning strategy that preserves spatial modeling capabilities while efficiently incorporating event-guided temporal information. We incorporate input-output normalization techniques inspired by recent advances in diffusion modeling to enhance training stability across varying noise levels. To improve generalization, we construct a comprehensive dataset combining both real and simulated event data across diverse scenarios. Extensive experiments on both real and simulated datasets demonstrate that EGVD significantly outperforms existing methods in handling large motion and challenging lighting conditions, achieving substantial improvements in perceptual quality metrics (27.4% better LPIPS on Prophesee and 24.1% on BSRGB) while maintaining competitive fidelity measures. Code and datasets available at: https://github.com/OpenImagingLab/EGVD.

视频插帧扩散模型事件相机大运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。