arXiv:2507.03765cs.CV2025-07AAAI被引 12

融合帧与事件数据,用轻量网络提升分割精度并降耗。

Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach

  • 双分支网络:脉冲神经网络处理事件,人工神经网络处理帧。
  • 在三个数据集上达顶尖精度,且能耗降低65%。
  • 适合低功耗实时视觉系统开发者参考。

事件相机因其高时间分辨率等优势被引入图像语义分割领域。然而,现有基于事件的分割方法未能充分挖掘帧与事件之间的互补信息,导致训练策略复杂且计算开销大。为此,我们提出一种高效的混合框架,包含用于事件的脉冲神经网络分支和用于帧的人工神经网络分支。设计了三个专用模块以促进两分支交互:自适应时间加权注入器(ATW)动态融合事件的时间特征到帧特征中,提升分割精度;事件驱动稀疏注入器(EDS)有效结合稀疏事件数据与丰富帧特征,实现精准时空对齐;通道选择融合模块(CSF)有选择性地融合特征以优化性能。实验表明,该框架在DDD17-Seg、DSEC-Semantic和M3ED-Semantic数据集上均达到领先精度,且显著降低能耗,在DSEC-Semantic数据集上减少65%能耗。

原文摘要 · Abstract (English)

Event cameras have recently been introduced into image semantic segmentation, owing to their high temporal resolution and other advantageous properties. However, existing event-based semantic segmentation methods often fail to fully exploit the complementary information provided by frames and events, resulting in complex training strategies and increased computational costs. To address these challenges, we propose an efficient hybrid framework for image semantic segmentation, comprising a Spiking Neural Network branch for events and an Artificial Neural Network branch for frames. Specifically, we introduce three specialized modules to facilitate the interaction between these two branches: the Adaptive Temporal Weighting (ATW) Injector, the Event-Driven Sparse (EDS) Injector, and the Channel Selection Fusion (CSF) module. The ATW Injector dynamically integrates temporal features from event data into frame features, enhancing segmentation accuracy by leveraging critical dynamic temporal information. The EDS Injector effectively combines sparse event data with rich frame features, ensuring precise temporal and spatial information alignment. The CSF module selectively merges these features to optimize segmentation performance. Experimental results demonstrate that our framework not only achieves state-of-the-art accuracy across the DDD17-Seg, DSEC-Semantic, and M3ED-Semantic datasets but also significantly reduces energy consumption, achieving a 65\% reduction on the DSEC-Semantic dataset.

事件相机语义分割低功耗多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。