用跨模态对齐提升脑电视觉解码精度
Neural-MCRL: Neural Multimodal Contrastive Representation Learning for EEG-based Visual Decoding
- 通过语义桥接与交叉注意力实现多模态对齐
- 在多个数据集上准确率显著优于现有方法
- 适合脑机接口与神经康复研究者使用
从脑电图(EEG)中解码神经视觉表征对于推进脑机接口(BMI)和神经感官康复具有重要意义。尽管多模态对比表示学习(MCRL)在神经解码中展现潜力,但现有方法常忽视模态内语义一致性与完整性,缺乏模态间有效语义对齐,限制了对复杂视觉神经反应的捕捉。本文提出Neural-MCRL框架,通过语义桥接与交叉注意力机制实现多模态对齐,同时保证模态内完整性与跨模态一致性。该框架引入具备谱-时序自适应能力的神经编码器(NESTA),可自适应捕获频谱模式并学习个体特异性变换。实验表明,相比最先进方法,本框架在视觉解码准确率与模型泛化能力上均有显著提升,推动了基于EEG的神经视觉表征解码发展。代码将公开于:https://github.com/NZWANG/Neural-MCRL。
原文摘要 · Abstract (English)
Decoding neural visual representations from electroencephalogram (EEG)-based brain activity is crucial for advancing brain-machine interfaces (BMI) and has transformative potential for neural sensory rehabilitation. While multimodal contrastive representation learning (MCRL) has shown promise in neural decoding, existing methods often overlook semantic consistency and completeness within modalities and lack effective semantic alignment across modalities. This limits their ability to capture the complex representations of visual neural responses. We propose Neural-MCRL, a novel framework that achieves multimodal alignment through semantic bridging and cross-attention mechanisms, while ensuring completeness within modalities and consistency across modalities. Our framework also features the Neural Encoder with Spectral-Temporal Adaptation (NESTA), a EEG encoder that adaptively captures spectral patterns and learns subject-specific transformations. Experimental results demonstrate significant improvements in visual decoding accuracy and model generalization compared to state-of-the-art methods, advancing the field of EEG-based neural visual representation decoding in BMI. Codes will be available at: https://github.com/NZWANG/Neural-MCRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。