arXiv:2608.01355cs.CV2026-08

通过融合多路视觉教师评分,提升脑电图像检索准确率

CORTIVA: Candidate-Score Fusion of Complementary Visual Teachers for EEG- and MEG-to-Image Retrieval

论文配图:CORTIVA: Candidate-Score Fusion of Complementary Visual Teachers for EEG- and MEG-to-Image Retrieval
图 1 · 摘自论文原文
  • 保留多路视觉教师独立评分,最后融合温度归一化得分
  • 在THINGS-EEG2上达73.5% Top-1,比基线高10.3个百分点
  • 方法简单可复现,适合脑机接口与神经解码研究者

从非侵入性脑活动解码视觉体验是神经科学与脑机接口的核心。功能磁共振(fMRI)虽空间分辨率高,但时间延迟大、采集负担重,难以实现时序解析解码。脑电图(EEG)和脑磁图(MEG)具备毫秒级时间分辨率,使图像检索成为可能:从一次神经反应和固定候选集里识别出观看的图像。基于预训练视觉表征的对比对齐可实现零样本检索,但多数系统将异构视觉监督提前合并为单一嵌入,导致所有候选排序共享同一相似性几何结构,并丢失编码器间的差异性信息。本文提出CORTIVA,一种候选评分融合框架,保留互补证据。三条解码路径分别对齐异构视觉目标,独立评分相同候选集,仅在排序前融合其温度缩放后的得分向量。在200分类的THINGS-EEG2基准上,CORTIVA在十名参与者中达到73.5% Top-1和95.3% Top-5,优于最强基线10.3和5.4个百分点。采用模态特异性神经编码器时,同一融合原则在THINGS-MEG上达42.4% Top-1。匹配路径移除重训及四种权重控制实验表明,性能提升源于互补路径得分整合,且在均匀加权下仍有效,无需专用权重规则。独立DINOv2分析进一步复现局部错误邻域与后验神经-视觉对应关系。这些结果确立候选评分融合为神经图像检索中替代嵌入级合并的简单而可验证的方案。

原文摘要 · Abstract (English)

Decoding visual experience from non-invasive brain activity is central to neuroscience and brain-computer interfaces. Functional magnetic resonance imaging (fMRI) offers fine spatial detail, but its slow hemodynamics and burdensome acquisition limit temporally resolved decoding. Electroencephalography (EEG) and magnetoencephalography (MEG) provide millisecond resolution, making image retrieval compelling: identify the viewed image from one neural response and a fixed candidate bank. Contrastive alignment to pretrained visual representations enables zero-shot retrieval from EEG and MEG, but most systems collapse heterogeneous visual supervision into a single embedding before ranking. This early consolidation imposes one similarity geometry on every candidate order and removes encoder-specific disagreements from the final ranking. We propose CORTIVA, a candidate-score fusion framework that preserves this complementary evidence. Three decoding routes are aligned to heterogeneous visual targets, score the same indexed candidates independently, and combine only their temperature-scaled score vectors before ranking. On the 200-way THINGS-EEG2 benchmark, CORTIVA reaches 73.5% Top-1 and 95.3% Top-5 across ten participants, exceeding the strongest reported baseline by 10.3 and 5.4 percentage points. With a modality-specific neural encoder, the same fusion principle reaches 42.4% Top-1 on THINGS-MEG. Matched route-removal retraining and four weight controls demonstrate that CORTIVA's gain arises from integrating complementary route scores and persists with uniform weighting, without requiring a specialized weighting rule. Independent DINOv2 analyses further reproduce the local error neighborhoods and posterior neural-visual correspondence. These results establish candidate-score fusion as a simple and testable alternative to embedding-level consolidation for neural image retrieval.

脑机接口图像检索EEG/MEG多路融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。