用事件相机与图像融合,精准预测视频插帧的运动轨迹。
Event-Based Video Frame Interpolation With Cross-Modal Asymmetric Bidirectional Motion Fields
- 结合事件与图像特征,直接估计非对称双向运动场。
- 在ERF-X170FPS数据集上,峰值信噪比提升至41.23dB。
- 适合需要高动态场景插帧的视觉系统开发者。
视频帧插值(VFI)旨在生成连续输入帧之间的中间帧。由于事件相机是生物启发式传感器,能以微秒级时间分辨率仅编码亮度变化,已有研究利用事件相机提升VFI性能。然而,现有方法仅使用事件或近似估算双向帧间运动场,难以应对真实场景中的复杂运动。本文提出一种新型基于事件的VFI框架,实现跨模态非对称双向运动场估计。具体而言,所提出的EIF-BiOFNet充分利用事件与图像的各自优势,无需近似直接估计帧间运动场。此外,设计了基于交互注意力的帧合成网络,有效融合基于光流变形和基于重建的特征。最后,构建大规模事件驱动的VFI数据集ERF-X170FPS,具备高帧率、极端运动和动态纹理特性,弥补先前数据集的局限。大量实验验证,本方法在多个数据集上显著优于现有最优VFI方法,峰值信噪比达41.23dB。
原文摘要 · Abstract (English)
Video Frame Interpolation (VFI) aims to generate intermediate video frames between consecutive input frames. Since the event cameras are bio-inspired sensors that only encode brightness changes with a micro-second temporal resolution, several works utilized the event camera to enhance the performance of VFI. However, existing methods estimate bidirectional inter-frame motion fields with only events or approximations, which can not consider the complex motion in real-world scenarios. In this paper, we propose a novel event-based VFI framework with cross-modal asymmetric bidirectional motion field estimation. In detail, our EIF-BiOFNet utilizes each valuable characteristic of the events and images for direct estimation of inter-frame motion fields without any approximation methods. Moreover, we develop an interactive attention-based frame synthesis network to efficiently leverage the complementary warping-based and synthesis-based features. Finally, we build a large-scale event-based VFI dataset, ERF-X170FPS, with a high frame rate, extreme motion, and dynamic textures to overcome the limitations of previous event-based VFI datasets. Extensive experimental results validate that our method shows significant performance improvement over the state-of-the-art VFI methods on various datasets. Our project pages are available at: https://github.com/intelpro/CBMNet
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。