针对脑电图像解码中跨被试差异导致的性能下降,提出统一伪特征编码框架提升跨模态对齐。
SUP-MCRL: Subject-aware Unified Pseudo-feature Coded Multimodal Contrastive Representation Learning for EEG Visual Decoding

- 引入语义感知视觉编码器与自适应跨被试增强机制,实现对注意力和个体差异的建模。
- 在THINGS-EEG数据集上零样本测试达到24.0%/52.9%(LOSO)Top-1/Top-5准确率,显著优于现有方法。
- 适合关注脑机接口跨被试泛化、多模态表示学习的研究者。
非侵入式脑机接口在从实验室刺激转向真实自然图像时性能显著下降。原因在于传统多模态对比表示学习模型仅优化几何距离对齐,忽视神经表征中的语义一致性与被试间差异及选择性注意。这导致模型易产生虚假的零样本匹配。为此,本文提出SUP-MCRL,一个融合三项协同机制的统一框架:(1) 语义实体感知视觉编码器(SAVE),无需预训练显著性模型即可学习空间注意力以提取语义内容;(2) 统一脑电信号增强器(UEE),采用多尺度空洞卷积与频带间注意力实现自适应跨被试鲁棒性;(3) 基于原型的渐进式增强器(PPA),通过EMA更新的伪特征池防止表示崩溃。在THINGS-EEG数据集上的零样本实验中,内被试准确率达66.0%/91.9%(Top-1/Top-5),LOSO准确率为24.0%/52.9%,显著超越现有方法,验证了结构化对齐监督在克服跨模态解码局限中的关键作用。代码已开源。
原文摘要 · Abstract (English)
Non-invasive brain-computer interfaces exhibit significant performance degradation when moving from controlled laboratory stimuli to real-world natural images. This degradation occurs because conventional multimodal contrastive representation learning models focus exclusively on optimizing geometric distance alignment, thereby failing to account for semantic consistency and inter-subject variability in neural representation and selective attention. As a result, these models are prone to producing spurious zero-shot matches. To address these limitations, we propose SUP-MCRL, a unified framework integrating three collaborative mechanisms: (1) a Semantic-entity Aware Visual Encoder (SAVE) that learns spatial attention to extract semantic content without relying on pre-trained saliency models; (2) a Unified EEG Enhancer (UEE) that employs multi-scale atrous convolutions and inter-band attention for adaptive cross-subject robustness; and (3) a Prototype-based Progressive Augmenter (PPA) that maintains an EMA-updated pseudo-feature pool to prevent representation collapse. Zero-shot experiments on the THINGS-EEG achieve 66.0%/91.9% (Top-1/Top-5) intra-subject and 24.0%/52.9% LOSO accuracy, significantly surpassing state-of-the-art methods and demonstrating that structured alignment supervision is key to overcoming the limitations of cross-modal decoding. Code is available at https://github.com/NZWANG/SUP-MCRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。