arXiv:2409.11953cs.CV2024-09被引 10

融合图像与事件流,实现高速场景下稳定点追踪

Tracking Any Point with Frame-Event Fusion Network at High Frame Rate

  • 用事件数据引导图像生成过程,融合多模态信息
  • 在EDS数据集上特征年龄预期降低24%
  • 适合高帧率、动态复杂的视觉追踪任务

基于图像帧的任意点追踪受帧率限制,在高速场景中易出现不稳定,且在真实应用中泛化能力有限。为此,我们提出一种图像-事件融合点追踪器FE-TAP,将图像帧的上下文信息与事件的高时间分辨率相结合,实现在多种挑战性条件下高帧率、鲁棒的点追踪。具体而言,设计了进化融合模块(EvoFusion),以事件为指导建模图像生成过程,有效整合双模态在不同频率下的信息。为获得更平滑的点轨迹,采用基于Transformer的迭代优化策略,逐步更新点轨迹与特征。大量实验表明,该方法优于当前最先进方法,尤其在EDS数据集上预期特征年龄提升24%。最后,通过自研高分辨率图像-事件同步设备,在真实驾驶场景中定性验证了算法鲁棒性。源代码将公开于https://github.com/ljx1002/FE-TAP。

原文摘要 · Abstract (English)

Tracking any point based on image frames is constrained by frame rates, leading to instability in high-speed scenarios and limited generalization in real-world applications. To overcome these limitations, we propose an image-event fusion point tracker, FE-TAP, which combines the contextual information from image frames with the high temporal resolution of events, achieving high frame rate and robust point tracking under various challenging conditions. Specifically, we designed an Evolution Fusion module (EvoFusion) to model the image generation process guided by events. This module can effectively integrate valuable information from both modalities operating at different frequencies. To achieve smoother point trajectories, we employed a transformer-based refinement strategy that updates the point's trajectories and features iteratively. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, particularly improving expected feature age by 24$\%$ on EDS datasets. Finally, we qualitatively validated the robustness of our algorithm in real driving scenarios using our custom-designed high-resolution image-event synchronization device. Our source code will be released at https://github.com/ljx1002/FE-TAP.

点追踪事件相机多模态融合高帧率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。