让脑电与视觉模型中间层对齐,提升非侵入式脑机接口的图像解码准确率。
Aligning What EEG Can See: Structural Representations for Brain-Vision Matching
- 将脑电信号与视觉模型中间层对齐,减少跨模态信息错配。
- 在THINGS-EEG数据集上实现84.6%零样本解码准确率,提升21.4%。
- 适用于多种脑电基线模型,通用性强,适合脑机接口研究者。
从脑电图(EEG)进行视觉解码已成为非侵入式脑机接口中极具前景的方向。现有方法主要将脑电信号与深度视觉模型最后一层语义嵌入对齐,但高度抽象的嵌入导致严重的跨模态信息错配。本文提出「神经可见性」概念,并设计EEG可见层选择策略,使脑电信号与视觉模型中间层对齐,以最小化错配。此外,为适应人类视觉处理的多阶段特性,提出分层互补融合(HCF)框架,联合整合不同层次的视觉表示。大量实验表明,该方法达到当前最佳性能,在THINGS-EEG数据集上实现84.6%的零样本视觉解码准确率(+21.4%),且在多种EEG基线上性能提升最高达129.8%,展现出优异的泛化能力。
原文摘要 · Abstract (English)
Visual decoding from electroencephalography (EEG) has emerged as a highly promising avenue for non-invasive brain-computer interfaces (BCIs). Existing EEG-based decoding methods predominantly align brain signals with the final-layer semantic embeddings of deep visual models. However, relying on these highly abstracted embeddings inevitably leads to severe cross-modal information mismatch. In this work, we introduce the concept of Neural Visibility and accordingly propose the EEG-Visible Layer Selection Strategy, aligning EEG signals with intermediate visual layers to minimize this mismatch. Furthermore, to accommodate the multi-stage nature of human visual processing, we propose a novel Hierarchically Complementary Fusion (HCF) framework that jointly integrates visual representations from different hierarchical levels. Extensive experiments demonstrate that our method achieves state-of-the-art performance, reaching an 84.6% accuracy (+21.4%) on zero-shot visual decoding on the THINGS-EEG dataset. Moreover, our method achieves up to a 129.8% performance gain across diverse EEG baselines, demonstrating its robust generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。