用脉冲神经网络提升音视频融合特征区分能力
Spiking Neural Network Feature Discrimination Boosts Modality Fusion
- 设计脉冲神经网络实现音视频模态的特征区分
- 在多个数据集上达到优于传统模型的分类准确率
- 适合对能效敏感的多模态实时系统研究者
特征区分是神经网络设计中的关键环节,直接影响分类能力和跨数据集泛化性能。高质量特征表示需具备高类内可分性,是当前最具挑战性的研究方向之一。传统深度神经网络(DNN)依赖复杂变换和深层结构,通常需数天训练且能耗巨大。脉冲神经网络(SNN)因其捕捉时空依赖的能力,成为处理多模态复杂任务的有力替代方案。本文提出一种面向多模态学习的SNN特征区分方法,聚焦音视频数据。采用深脉冲残差网络处理视觉模态,用更简洁高效的脉冲网络处理听觉模态,并通过脉冲多层感知机实现模态融合。在多种分类任务中验证了该方法的有效性。据我们所知,这是首个系统研究SNN中特征区分的工作。
原文摘要 · Abstract (English)
Feature discrimination is a crucial aspect of neural network design, as it directly impacts the network's ability to distinguish between classes and generalize across diverse datasets. The accomplishment of achieving high-quality feature representations ensures high intra-class separability and poses one of the most challenging research directions. While conventional deep neural networks (DNNs) rely on complex transformations and very deep networks to come up with meaningful feature representations, they usually require days of training and consume significant energy amounts. To this end, spiking neural networks (SNNs) offer a promising alternative. SNN's ability to capture temporal and spatial dependencies renders them particularly suitable for complex tasks, where multi-modal data are required. In this paper, we propose a feature discrimination approach for multi-modal learning with SNNs, focusing on audio-visual data. We employ deep spiking residual learning for visual modality processing and a simpler yet efficient spiking network for auditory modality processing. Lastly, we deploy a spiking multilayer perceptron for modality fusion. We present our findings and evaluate our approach against similar works in the field of classification challenges. To the best of our knowledge, this is the first work investigating feature discrimination in SNNs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。