arXiv:2605.23996cs.CVeess.IV2026-05被引 1

用脑电波还原人看的图像,准确率超86%

Brain-to-Image Retrieval and Reconstruction via Multimodal EEG Alignment

论文配图:Brain-to-Image Retrieval and Reconstruction via Multimodal EEG Alignment
图 1 · 摘自论文原文
  • 通过多层级模糊+生物启发特征提升脑电到图像检索精度
  • 单人10次实验平均顶1准确率达86.3%,顶5达98.5%
  • 结合多模态对齐与生成模型,可重建接近真实视觉内容的图像

我们提出一种脑电到图像系统,从自然图像观看时记录的脑电(EEG)信号中解码视觉刺激。该系统解决两个任务:(1) 脑电到图像检索,给定一段脑电信号,在200张候选图中排序出正确刺激图像;(2) 脑电到图像重建,生成与感知内容一致的图像。检索方面,采用多层级模糊方法并结合生物启发的EVNet特征,使用InfoNCE损失训练。在单个受试者上进行10次随机种子评估,检索模型在最终训练轮次的平均Top-1准确率为86.30%,Top-5准确率为98.55%。重建方面,提出CognitionCapturerPro,将脑电表示对齐至多模态CLIP嵌入(包括图像、文本、深度、边缘嵌入),并使用以IP-Adapter为条件的SDXL-Turbo生成图像。平均10次种子测试,重建模型在ViT-H-14下取得0.903的CLIP分数,ViT-L/14下为0.870,SSIM为0.409。结果表明,利用现代多模态对齐与生成建模技术,从脑电中解码丰富视觉表征是可行的。

原文摘要 · Abstract (English)

We present a brain-to-image system that decodes visual stimuli from EEG signals recorded during natural image viewing. Our system addresses two tasks: (1) EEG-to-image retrieval, which ranks the correct stimulus image among 200 candidates given an EEG segment, and (2) EEG-to-image reconstruction, which generates an image consistent with the perceived stimulus. For retrieval, we implement a multi-level blurring approach improved with biologically inspired EVNet features and trained with the InfoNCE loss. Evaluated over 10 random seeds for a single subject, the retrieval model achieves a mean final-epoch Top-1 accuracy of 86.30% and Top-5 accuracy of 98.55%. For reconstruction, we implement CognitionCapturerPro, which aligns EEG representations to multi-modal CLIP embeddings, including image, text, depth, and edge embeddings, and synthesizes images with SDXL-Turbo conditioned via IP-Adapter. Averaged over 10 seeds, the reconstruction model achieves a CLIP score of 0.903 using ViT-H-14, a CLIP score of 0.870 using ViT-L/14, and an SSIM of 0.409. These results demonstrate the feasibility of decoding rich visual representations from EEG signals using modern multi-modal alignment and generative modeling techniques.

脑机接口图像重建多模态对齐EEG解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。