arXiv:2606.25718cs.CV2026-06

通过多视角建模提升脑电视觉解码精度,实现更准确的零样本图像识别。

What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

论文配图:What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment
图 1 · 摘自论文原文
  • 构建三视图脑电信号模型:时序、频谱与电极关联联合建模
  • 在同被试设置下达54.8%准确率,跨被试达15.3%,性能领先
  • 首次系统评估跨会话解码,验证方法泛化能力,适合脑机接口研究者

零样本脑电(EEG)视觉解码旨在从非侵入性神经信号中推断视觉语义,但受限于脑电信噪比低、非平稳性和空间分辨率差。现有方法多依赖整体脑电嵌入,掩盖了视觉感知背后的时序、频谱和空间结构。本文提出统一的多视角脑电表征学习框架,构建一个联合建模三种互补视图的编码器:输入条件的状态空间时序动态、可学习的小波基频谱分解以实现样本自适应频率建模,以及注意力调制的图学习以捕捉电极间结构化交互。最终的多视图脑电嵌入通过对比学习与特定脑电正则化,在共享语义空间中对齐预训练视觉表示,支持200类零样本图像分类。在THINGS-EEG基准测试中,该方法在同被试设置下达到54.8% Top-1和85.6% Top-5准确率,跨被试设置下为15.3% Top-1和45.4% Top-5。此外,首次系统评估跨会话解码,取得40.8% Top-1和78.0% Top-5准确率。结果表明,显式建模多视角神经结构能显著提升脑电-视觉对齐效果与泛化能力。

原文摘要 · Abstract (English)

Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited spatial resolution of EEG. Existing EEG-vision alignment methods often rely on holistic EEG embeddings, which can obscure the complementary temporal, spectral, and spatial structure underlying visual perception. We introduce a unified multiview EEG representation learning framework for aligning brain responses with visual semantic embeddings. Our method builds an EEG encoder that jointly models three complementary views: input-conditioned state-space temporal dynamics, learnable wavelet-based spectral decomposition for sample-adaptive frequency modeling, and attention-modulated graph learning for structured electrode interactions. The resulting multiview EEG embeddings are fused and aligned with pretrained visual representations in a shared semantic space using contrastive learning with EEG-specific regularization, enabling 200-way zero-shot visual classification. Experiments on THINGS-EEG benchmark show that our method achieves state-of-the-art performance, with 54.8% Top-1 and 85.6% Top-5 accuracy in the within-subject setting and 15.3% Top-1 and 45.4% Top-5 accuracy in the cross-subject setting. We further present the first systematic cross-session EEG-image decoding evaluation, achieving 40.8% Top-1 and 78.0% Top-5 accuracy. These results suggest that explicitly modeling multiview neural structure improves both semantic alignment and generalization in EEG-based visual decoding.

脑机接口视觉解码多视图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。