arXiv:2410.15078eess.ASeess.SP2024-10中稿 · the International …

通过跨模态融合提升语音与脑电匹配分类准确率

Independent Feature Enhanced Crossmodal Fusion for Match-Mismatch Classification of Speech Stimulus and EEG Response

  • 双分支结构结合交叉注意力,捕捉语音与脑电关系
  • 融合交互特征与独立特征,提升分类精度至91.2%
  • 引入因果掩码,适配语音-脑电时间延迟,适合神经解码研究者

在听觉注意解码中,准确分类语音刺激与对应脑电响应的匹配与不匹配关系至关重要。现有方法通常采用两个独立网络分别编码语音与脑电信号,忽略了两模态间的关联性。本文提出独立特征增强的跨模态融合模型(IFE-CF),通过融合语音与脑电信号实现听觉脑电解码。具体包括:基于交叉注意力机制的双分支跨模态编码器,用于捕捉两模态间动态关系;多通道融合模块,聚合来自编码器的交互特征与各模态的独立特征;以及预测模块输出匹配结果。此外,引入因果掩码以建模语音-脑电对的时间延迟,进一步增强特征表示。实验表明,本方法在2023年听觉脑电解码挑战赛基线基础上显著提升分类准确率,达到91.2%。

原文摘要 · Abstract (English)

It is crucial for auditory attention decoding to classify matched and mismatched speech stimuli with corresponding EEG responses by exploring their relationship. However, existing methods often adopt two independent networks to encode speech stimulus and EEG response, which neglect the relationship between these signals from the two modalities. In this paper, we propose an independent feature enhanced crossmodal fusion model (IFE-CF) for match-mismatch classification, which leverages the fusion feature of the speech stimulus and the EEG response to achieve auditory EEG decoding. Specifically, our IFE-CF contains a crossmodal encoder to encode the speech stimulus and the EEG response with a two-branch structure connected via crossmodal attention mechanism in the encoding process, a multi-channel fusion module to fuse features of two modalities by aggregating the interaction feature obtained from the crossmodal encoder and the independent feature obtained from the speech stimulus and EEG response, and a predictor to give the matching result. In addition, the causal mask is introduced to consider the time delay of the speech-EEG pair in the crossmodal encoder, which further enhances the feature representation for match-mismatch classification. Experiments demonstrate our method's effectiveness with better classification accuracy, as compared with the baseline of the Auditory EEG Decoding Challenge 2023.

跨模态融合脑电解码语音识别注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。