arXiv:2411.18658cs.CV2024-11被引 7

首个直接训练的帧事件混合神经架构,兼顾精度与低功耗

HDI-Former: Hybrid Dynamic Interaction ANN-SNN Transformer for Object Detection Using Frames and Events

  • 设计帧事件双流动态交互机制,提升跨模态信息融合
  • 在DSEC数据集上比肩纯ANN性能,能耗降低10.57倍
  • 适合做低功耗智能视觉系统的研究与开发人员

融合帧图像与事件流可有效应对复杂场景下的目标检测挑战。然而,现有方法多采用独立的人工神经网络(ANN)分支,限制了跨模态信息交互,且难以从事件流中高效提取时序特征并实现低功耗运行。为此,本文提出HDI-Former——首个直接训练的帧事件混合型ANN-SNN Transformer架构,兼顾高精度与低功耗。技术上,首先提出语义增强的自注意力机制,强化ANN分支中图像编码标记间的关联性;其次设计脉冲Swin Transformer分支,以低功耗建模事件流中的时序线索;最后引入类生物启发的动态交互机制,实现ANN与SNN子网络间的跨模态信息交互。实验表明,HDI-Former显著优于11种先进方法及4个基线模型。其SNN分支在保持与同结构ANN相当性能的同时,在DSEC-Detection数据集上能耗降低10.57倍。开源代码已提供于补充材料。

原文摘要 · Abstract (English)

Combining the complementary benefits of frames and events has been widely used for object detection in challenging scenarios. However, most object detection methods use two independent Artificial Neural Network (ANN) branches, limiting cross-modality information interaction across the two visual streams and encountering challenges in extracting temporal cues from event streams with low power consumption. To address these challenges, we propose HDI-Former, a Hybrid Dynamic Interaction ANN-SNN Transformer, marking the first trial to design a directly trained hybrid ANN-SNN architecture for high-accuracy and energy-efficient object detection using frames and events. Technically, we first present a novel semantic-enhanced self-attention mechanism that strengthens the correlation between image encoding tokens within the ANN Transformer branch for better performance. Then, we design a Spiking Swin Transformer branch to model temporal cues from event streams with low power consumption. Finally, we propose a bio-inspired dynamic interaction mechanism between ANN and SNN sub-networks for cross-modality information interaction. The results demonstrate that our HDI-Former outperforms eleven state-of-the-art methods and our four baselines by a large margin. Our SNN branch also shows comparable performance to the ANN with the same architecture while consuming 10.57$\times$ less energy on the DSEC-Detection dataset. Our open-source code is available in the supplementary material.

目标检测事件相机脉冲神经网络低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。