arXiv:2505.07715cs.CVcs.AI2025-05ICML被引 13

用脉冲神经网络提升事件相机目标检测,更准更省电。

Hybrid Spiking Vision Transformer for Object Detection with Event Cameras

  • 融合空间与时间模块,捕捉事件数据的时空特征。
  • 在多个数据集上参数更少,检测准确率显著提升。
  • 适合低功耗、实时目标检测场景的研究与应用。

基于事件的物体检测因其高时间分辨率、宽动态范围和异步地址-事件表示等优势而受到越来越多关注。利用这些优势,脉冲神经网络(SNN)成为一种有前景的方法,具有低能耗和丰富的时空动态特性。为进一步提升基于事件的物体检测性能,本文提出一种新型混合脉冲视觉变换器(HsVT)模型。该模型结合空间特征提取模块以捕获局部与全局特征,以及时间特征提取模块以建模事件序列中的时间依赖性和长期模式。这种组合使HsVT能够有效捕捉时空特征,增强其处理复杂事件感知任务的能力。为支持该领域研究,我们构建并公开发布“跌倒检测数据集”作为基准,该数据集由事件相机采集,确保面部隐私保护,并因事件表示格式减少内存占用。我们在GEN1和跌倒检测数据集上对HsVT模型进行了评估,涵盖多种模型规模。实验结果表明,相较于现有方法,HsVT在事件检测任务中以更少参数实现了显著性能提升。

原文摘要 · Abstract (English)

Event-based object detection has gained increasing attention due to its advantages such as high temporal resolution, wide dynamic range, and asynchronous address-event representation. Leveraging these advantages, Spiking Neural Networks (SNNs) have emerged as a promising approach, offering low energy consumption and rich spatiotemporal dynamics. To further enhance the performance of event-based object detection, this study proposes a novel hybrid spike vision Transformer (HsVT) model. The HsVT model integrates a spatial feature extraction module to capture local and global features, and a temporal feature extraction module to model time dependencies and long-term patterns in event sequences. This combination enables HsVT to capture spatiotemporal features, improving its capability to handle complex event-based object detection tasks. To support research in this area, we developed and publicly released The Fall Detection Dataset as a benchmark for event-based object detection tasks. This dataset, captured using an event-based camera, ensures facial privacy protection and reduces memory usage due to the event representation format. We evaluated the HsVT model on GEN1 and Fall Detection datasets across various model sizes. Experimental results demonstrate that HsVT achieves significant performance improvements in event detection with fewer parameters.

事件相机脉冲神经网络目标检测低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。