arXiv:2410.15689cs.CVcs.LG2024-10被引 17

构建新数据集与跨模态注意力模型,提升脉冲神经网络的时空学习能力。

Enhancing SNN-based Spatio-Temporal Learning: A Benchmark Dataset and Cross-Modality Attention Model

  • 提出DVS-SLR数据集,具更高时间相关性与多场景覆盖
  • 跨模态注意力模型使脉冲网络同时学习时空注意力并增强融合效果
  • 适合研究脉冲神经网络、多模态融合与类脑计算的学者

脉冲神经网络(SNN)因其低功耗、类脑架构和时空表征能力受到广泛关注。然而,现有类脑数据集普遍缺乏强时间相关性,限制了SNN对时空特性的发挥。同时,事件流与图像帧的融合可提供更全面的视觉时空信息,但基于SNN的跨模态融合仍不充分。本文提出一个名为DVS-SLR的新类脑数据集,其在时间相关性、规模和场景多样性上均优于现有数据集,并包含对应帧数据,支持开发双模态融合方法。基于该数据集,我们设计了一种跨模态注意力(CMA)融合模型,使SNN能从事件与帧模态中分别学习时空注意力分数,并动态分配以增强模态协同。实验表明,该方法显著提升识别准确率,并在多样场景下保持鲁棒性。

原文摘要 · Abstract (English)

Spiking Neural Networks (SNNs), renowned for their low power consumption, brain-inspired architecture, and spatio-temporal representation capabilities, have garnered considerable attention in recent years. Similar to Artificial Neural Networks (ANNs), high-quality benchmark datasets are of great importance to the advances of SNNs. However, our analysis indicates that many prevalent neuromorphic datasets lack strong temporal correlation, preventing SNNs from fully exploiting their spatio-temporal representation capabilities. Meanwhile, the integration of event and frame modalities offers more comprehensive visual spatio-temporal information. Yet, the SNN-based cross-modality fusion remains underexplored. In this work, we present a neuromorphic dataset called DVS-SLR that can better exploit the inherent spatio-temporal properties of SNNs. Compared to existing datasets, it offers advantages in terms of higher temporal correlation, larger scale, and more varied scenarios. In addition, our neuromorphic dataset contains corresponding frame data, which can be used for developing SNN-based fusion methods. By virtue of the dual-modal feature of the dataset, we propose a Cross-Modality Attention (CMA) based fusion method. The CMA model efficiently utilizes the unique advantages of each modality, allowing for SNNs to learn both temporal and spatial attention scores from the spatio-temporal features of event and frame modalities, subsequently allocating these scores across modalities to enhance their synergy. Experimental results demonstrate that our method not only improves recognition accuracy but also ensures robustness across diverse scenarios.

脉冲神经网络跨模态融合类脑计算时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。