用图像引导事件数据融合,提升动态场景光流估计精度与速度
Spatially-guided Temporal Aggregation for Robust Event-RGB Optical Flow Estimation
- 以图像为引导,聚合高时序密度的事件数据
- 在DSEC-Flow上准确率领先,比纯事件模型提升10%
- 适合需要高速高精度光流的自动驾驶等场景
当前光流方法依赖帧(或RGB)数据的稳定外观建立时间对应关系。事件相机则提供高时序分辨率的运动线索,在挑战性场景中表现优异。两者互补特性表明融合帧与事件数据具有潜力。然而,多数跨模态方法仅简单堆叠信息,未能充分利用优势。本文提出一种新方法:以空间稠密的模态引导时序密集事件模态的聚合,实现有效跨模态融合。具体地,构建事件增强的帧表示,保留帧的丰富纹理和事件的基本结构;以该增强表示作为引导模态,利用事件捕捉时序密集的运动信息。由引导模态提取的鲁棒运动特征指导事件运动信息的聚合。为进一步提升融合效果,设计基于Transformer的模块,将稀疏事件运动特征与空间丰富的帧信息互补,并增强全局信息传播。此外,提出混合融合编码器,从双模态中提取全面的时空上下文特征。在MVSEC和DSEC-Flow数据集上的大量实验表明,本框架有效。利用帧与事件的互补优势,方法在DSEC-Flow上达到领先性能。相比纯事件模型,帧引导使准确率提升10%;且优于现有最佳融合方法,准确率提高4%,推理时间减少45%。
原文摘要 · Abstract (English)
Current optical flow methods exploit the stable appearance of frame (or RGB) data to establish robust correspondences across time. Event cameras, on the other hand, provide high-temporal-resolution motion cues and excel in challenging scenarios. These complementary characteristics underscore the potential of integrating frame and event data for optical flow estimation. However, most cross-modal approaches fail to fully utilize the complementary advantages, relying instead on simply stacking information. This study introduces a novel approach that uses a spatially dense modality to guide the aggregation of the temporally dense event modality, achieving effective cross-modal fusion. Specifically, we propose an event-enhanced frame representation that preserves the rich texture of frames and the basic structure of events. We use the enhanced representation as the guiding modality and employ events to capture temporally dense motion information. The robust motion features derived from the guiding modality direct the aggregation of motion information from events. To further enhance fusion, we propose a transformer-based module that complements sparse event motion features with spatially rich frame information and enhances global information propagation. Additionally, a mix-fusion encoder is designed to extract comprehensive spatiotemporal contextual features from both modalities. Extensive experiments on the MVSEC and DSEC-Flow datasets demonstrate the effectiveness of our framework. Leveraging the complementary strengths of frames and events, our method achieves leading performance on the DSEC-Flow dataset. Compared to the event-only model, frame guidance improves accuracy by 10\%. Furthermore, it outperforms the state-of-the-art fusion-based method with a 4\% accuracy gain and a 45\% reduction in inference time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。