用事件相机信息提升视频插帧质量,减少模糊和错位。
Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

- 用事件流生成时空对齐的运动与结构引导信号。
- 在真实与合成数据集上均超越现有最优方法。
- 无需重新训练,仅通过适配器融合事件信息。
潜在扩散模型近期通过生成输入图像间的中间帧,推动了视频帧插值的发展。然而,在处理大时间间隔和复杂运动时仍面临挑战,常导致运动模糊、结构失真和时间不一致。事件相机提供高时间分辨率的运动线索,非常适合填补这些间隙并提升插值质量。为利用这一优势而不需从零开始训练事件辅助模型,我们提出一种基于适配器的框架,以最小架构改动将事件导出的线索注入预训练的图像到视频扩散模型。具体地,方法利用图像扭曲事件(IWEs)和双向稀疏光流,在生成过程中提供空间与时间对齐的引导。通过将这些事件引导的结构与运动线索注入扩散过程,该方法显著减少了插值伪影,提升了重建保真度与时间连贯性。在真实与合成基准上的实验结果表明,本方法持续优于现有最先进方法。
原文摘要 · Abstract (English)
Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling large temporal gaps and complex motion remains challenging, often resulting in motion blur, structural distortions, and temporal inconsistencies. Event cameras provide high-temporal-resolution motion cues that are well suited for bridging these gaps and improving interpolation quality. To exploit this advantage without training an event-assisted model from scratch, we propose an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes. Specifically, our method leverages Image Warped Events (IWEs) and bidirectional sparse optical flow to provide spatially and temporally aligned guidance during generation. By injecting these event-guided structural and motion cues into the diffusion process, our approach reduces interpolation artifacts and improves both reconstruction fidelity and temporal coherence. Experimental results on real and synthetic benchmarks show that our method consistently outperforms existing state-of-the-art approaches. The project page is at https://joseph-lin-tech.github.io/BridgeEventDiT-VFI/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。