通过锚点对齐神经与视觉表示,提升少次重复脑图检索精度。
Beyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval

- 以高重复中心为锚点,双向优化噪声脑信号与图像特征。
- 在仅1次或4次重复时,THINGS-EEG2上准确率提升5.7和9.3个百分点。
- 适用于真实场景下低重复刺激的脑机图像检索,降低用户负担。
从脑信号解码视觉信息可揭示神经表征,支持神经康复与梦解析。近期脑图检索方法表现良好,通常需对每张图像平均80次神经试验,导致延迟、成本上升及用户负担加重。当仅有一到几次重复时,检索准确率急剧下降。传统认为这是由查询噪声引起,因平均能抑制噪声、增强信号稳定性。然而我们发现非传递对齐模式:低重复查询信号与图像表征各自与高重复中心对齐,但彼此不直接对齐。这表明噪声并非唯一问题,画廊位置也影响检索。因此提出神经锚点检索框架(NEAR),将高重复中心作为锚点,从两侧逼近:去噪器将噪声查询拉向真实锚点,小型网络从图像预测伪锚点并拉动图像靠近。在四个涵盖EEG、MEG和fMRI的数据集上,NEAR在少次重复情形下持续提升检索性能。在THINGS-EEG2上,平均1次和4次重复时,200类Top-1准确率分别提升5.7和9.3个百分点。通过锚定神经与视觉表示,NEAR减少对重复采集的依赖,推动神经检索向实际应用迈进。
原文摘要 · Abstract (English)
Decoding visual information from brain signals probes neural representations and enables neuro-rehabilitation and dream decoding. Recent brain-to-image retrieval approaches have achieved promising performance, typically by averaging many (up to 80) neural trials per image, requiring repeated stimulus presentation that increases latency, cost, and user burden. When only one or a few repetitions are available, the retrieval accuracy drops sharply. This drop is commonly attributed to query noise because averaging suppresses noise and increases signal stability. However, we find a non-transitive alignment pattern: the low-repetition query signal and the image representation each align with the high-repetition center, but not directly with each other. This pattern shows that query noise is only part of the problem and that gallery placement also affects retrieval. We therefore propose a neural-anchor-based retrieval (NEAR) framework that treats the high-repetition center as an anchor and approaches it from both sides: a denoiser pulls the noisy query toward the true anchor, and a small network predicts each candidate's pseudo anchor from its image and pulls the image toward it. Across four datasets spanning EEG, MEG and fMRI, NEAR consistently improved retrieval in the few-repetition regime. On THINGS-EEG2, it improved 200-way Top-1 accuracy by 5.7 and 9.3 percentage points respectively, when averaging one and four repetitions. By anchoring neural and visual representations, NEAR reduces reliance on repeated acquisition and brings neural retrieval closer to real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。