arXiv:2507.21971cs.CV2025-07被引 1

融合事件与图像数据,提升复杂环境下的语义分割精度。

EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation

  • 通过多尺度建模和空间注意力优化稀疏事件流特征。
  • 在DDD17-Semantic和DSEC-Semantic上达到当前最优性能。
  • 适合关注低光照、高速运动场景视觉理解的研究者。

事件相机具有高动态范围和精细时间分辨率,有助于在复杂环境下实现鲁棒的场景理解。然而,该任务仍面临两大挑战:从稀疏且嘈杂的事件流中提取可靠特征,以及有效融合结构和表示差异显著的密集语义图像数据。为此,我们提出EIFNet,一种结合事件与帧输入优势的多模态融合网络。网络包含自适应事件特征精炼模块(AEFRM),通过多尺度活动建模与空间注意力提升事件表征;引入模态自适应重校准模块(MARM)和多头注意力门控融合模块(MGFM),利用注意力机制与门控融合策略实现跨模态特征对齐与整合。在DDD17-Semantic与DSEC-Semantic数据集上的实验表明,EIFNet取得当前最优性能,验证了其在事件驱动语义分割中的有效性。

原文摘要 · Abstract (English)

Event-based semantic segmentation explores the potential of event cameras, which offer high dynamic range and fine temporal resolution, to achieve robust scene understanding in challenging environments. Despite these advantages, the task remains difficult due to two main challenges: extracting reliable features from sparse and noisy event streams, and effectively fusing them with dense, semantically rich image data that differ in structure and representation. To address these issues, we propose EIFNet, a multi-modal fusion network that combines the strengths of both event and frame-based inputs. The network includes an Adaptive Event Feature Refinement Module (AEFRM), which improves event representations through multi-scale activity modeling and spatial attention. In addition, we introduce a Modality-Adaptive Recalibration Module (MARM) and a Multi-Head Attention Gated Fusion Module (MGFM), which align and integrate features across modalities using attention mechanisms and gated fusion strategies. Experiments on DDD17-Semantic and DSEC-Semantic datasets show that EIFNet achieves state-of-the-art performance, demonstrating its effectiveness in event-based semantic segmentation.

事件相机语义分割多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。