arXiv:2604.17782cs.CV2026-04被引 1

让脑电图与图像对齐更精准,因人而异构建视觉目标。

Subject-Aware Multi-Granularity Alignment for Zero-Shot EEG-to-Image Retrieval

论文配图:Subject-Aware Multi-Granularity Alignment for Zero-Shot EEG-to-Image Retrieval
图 1 · 摘自论文原文
  • 根据个体差异动态构建多粒度视觉监督信号
  • 在THINGS-EEG数据集上提升12%检索准确率
  • 适合开发个性化无创脑机接口的研究者

从脑电图(EEG)解码视觉内容对于理解神经视觉表征和开发非侵入式脑机接口具有重要意义。现有方法主要提升EEG表示学习与跨模态对齐,但将预训练视觉模型视为固定监督目标。然而,预训练视觉模型以层次化方式组织信息,不同深度编码互补的结构与语义信息,且与EEG最匹配的视觉粒度在不同受试者间存在差异。为此,我们提出主体感知的多粒度对齐(SAMGA),将视觉目标构建显式纳入EEG-视觉对齐过程。SAMGA从多个中间表示中构建自适应视觉监督,并通过全局粒度先验与受试者相关残差校准建模适配EEG的视觉粒度,实现主体感知训练与主体无关推理。基于此自适应目标,采用粗到精的对齐策略,先组织全局跨模态几何结构,再细化实例级检索判别能力。在THINGS-EEG数据集上,SAMGA在同主体评估下相比最强竞争方法提升8.7个百分点,在留一主体评估下提升12.0个百分点。结果表明,神经-视觉对齐性能不仅取决于神经表示的映射方式,还取决于定义其监督的视觉表示本身。

原文摘要 · Abstract (English)

Decoding visual content from electroencephalography (EEG) is important for understanding neural visual representations and developing non-invasive brain-computer interfaces. Existing approaches mainly improve EEG representation learning and cross-modal alignment while treating pretrained visual representations as fixed supervision targets. However, pretrained vision models organize information hierarchically, with different depths encoding complementary structural and semantic information, and the visual granularity most compatible with EEG may vary across subjects. To address this issue, we propose Subject-Aware Multi-Granularity Alignment (SAMGA), which makes visual-target construction an explicit part of EEG-visual alignment. SAMGA constructs adaptive visual supervision from multiple intermediate representations and models EEG-compatible visual granularity through a global granularity prior with subject-dependent residual calibration, enabling subject-aware training and subject-agnostic inference. Based on the resulting adaptive target, a coarse-to-fine alignment strategy first organizes global cross-modal geometry and then refines instance-level retrieval discrimination. On THINGS-EEG, SAMGA improves Top-1 retrieval accuracy over the strongest competing method by 8.7 percentage points under intra-subject evaluation and 12.0 percentage points under leave-one-subject-out evaluation. These results support a broader view of neural-visual alignment, in which performance depends not only on how neural representations are mapped, but also on what visual representations define their supervision.

脑机接口多模态对齐个性化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。