arXiv:2409.00552eess.AScs.CV2024-09被引 3

用多模态脉冲神经网络识别数字,准确率达98.43%

Digit Recognition using Multimodal Spiking Neural Networks

  • 融合视觉与听觉脉冲信号,在深层进行特征拼接
  • 多模态模型准确率98.43%,优于单一模态
  • 融合策略对融合深度不敏感,鲁棒性强

脉冲神经网络(SNN)是第三代类脑神经网络,通过模拟大脑信号传递方式处理数据。在计算机视觉领域,事件传感器的出现使基于事件的数据采集成为可能,其输出为随时间变化的脉冲序列,非常适合SNN处理。本文研究多模态输入在分类任务中的优势,采用视觉模态(Neuromorphic-MNIST,简称N-MNIST)和听觉模态(Spiking Heidelberg Digits,简称SHD)两个事件传感器生成的数据集,构建多模态SNN模型。实验表明,多模态SNN在融合视觉与听觉信息后,性能显著优于仅使用单一模态的SNN。此外,不同融合深度对结果影响较小,表明感官融合机制具有较强的鲁棒性。最终模型在联合N-MNIST与SHD数据集上达到98.43%的分类准确率。

原文摘要 · Abstract (English)

Spiking neural networks (SNNs) are the third generation of neural networks that are biologically inspired to process data in a fashion that emulates the exchange of signals in the brain. Within the Computer Vision community SNNs have garnered significant attention due in large part to the availability of event-based sensors that produce a spatially resolved spike train in response to changes in scene radiance. SNNs are used to process event-based data due to their neuromorphic nature. The proposed work examines the neuromorphic advantage of fusing multiple sensory inputs in classification tasks. Specifically we study the performance of a SNN in digit classification by passing in a visual modality branch (Neuromorphic-MNIST [N-MNIST]) and an auditory modality branch (Spiking Heidelberg Digits [SHD]) from datasets that were created using event-based sensors to generate a series of time-dependent events. It is observed that multi-modal SNNs outperform unimodal visual and unimodal auditory SNNs. Furthermore, it is observed that the process of sensory fusion is insensitive to the depth at which the visual and auditory branches are combined. This work achieves a 98.43% accuracy on the combined N-MNIST and SHD dataset using a multimodal SNN that concatenates the visual and auditory branches at a late depth.

脉冲神经网络多模态图像识别类脑计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。