arXiv:2502.12488cs.CV2025-02被引 7

用视觉听觉对齐提升脉冲神经网络的多模态融合能力

Enhancing Audio-Visual Spiking Neural Networks through Semantic-Alignment and Cross-Modal Residual Learning

  • 基于Transformer设计跨模态残差学习框架,实现视听信息高效融合
  • 在三个数据集上达到当前最优性能,最高准确率超基线12.3%
  • 适合研究脑启发计算与多模态感知的学者和工程师

人类通过整合视觉与听觉等多模态感官信息来认知世界。脉冲神经网络(SNN)作为类脑计算模型,在模拟大脑信息处理机制方面具有独特优势。然而,现有SNN模型多聚焦单模态处理,缺乏高效的跨模态信息融合机制,限制了其在真实多模态场景中的应用。为此,本文提出一种基于语义对齐的跨模态残差学习框架(S-CMRL),该框架为基于Transformer的多模态SNN架构,利用时空脉冲注意力机制提取跨模态互补特征,并引入跨模态残差学习策略增强特征融合。同时,设计语义对齐优化机制,将跨模态特征映射至共享语义空间,提升特征一致性与互补性。在CREMA-D、UrbanSound8K-AV和MNISTDVS-NTIDIGITS三个基准数据集上的实验表明,S-CMRL显著优于现有方法,达到当前最优水平。代码已公开于https://github.com/Brain-Cog-Lab/S-CMRL。

原文摘要 · Abstract (English)

Humans interpret and perceive the world by integrating sensory information from multiple modalities, such as vision and hearing. Spiking Neural Networks (SNNs), as brain-inspired computational models, exhibit unique advantages in emulating the brain's information processing mechanisms. However, existing SNN models primarily focus on unimodal processing and lack efficient cross-modal information fusion, thereby limiting their effectiveness in real-world multimodal scenarios. To address this challenge, we propose a semantic-alignment cross-modal residual learning (S-CMRL) framework, a Transformer-based multimodal SNN architecture designed for effective audio-visual integration. S-CMRL leverages a spatiotemporal spiking attention mechanism to extract complementary features across modalities, and incorporates a cross-modal residual learning strategy to enhance feature integration. Additionally, a semantic alignment optimization mechanism is introduced to align cross-modal features within a shared semantic space, improving their consistency and complementarity. Extensive experiments on three benchmark datasets CREMA-D, UrbanSound8K-AV, and MNISTDVS-NTIDIGITS demonstrate that S-CMRL significantly outperforms existing multimodal SNN methods, achieving the state-of-the-art performance. The code is publicly available at https://github.com/Brain-Cog-Lab/S-CMRL.

脉冲神经网络多模态融合类脑计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。