arXiv:2507.07157cs.CVcs.LG2025-07被引 2

用语义提示让脑电数据生成可解释的图像,突破传统直接重建瓶颈。

Interpretable EEG-to-Image Generation with Semantic Prompts

  • 通过大语言模型生成多层级语义描述,将脑电信号映射到这些描述上
  • 在EEGCVPR数据集上实现当前最佳视觉解码性能,生成图像与认知路径对齐
  • 适合脑机接口、神经科学解释性研究者使用,结果具可解释性

从脑信号中解码视觉体验为神经科学和可解释人工智能带来了新可能。尽管脑电图(EEG)具有高时间分辨率且易于获取,但其空间细节不足限制了图像重建效果。本研究摒弃直接的EEG到图像生成,转而将脑电信号与由大语言模型生成的多层次语义描述(从物体级到抽象主题)进行对齐。基于Transformer的EEG编码器通过对比学习将脑活动映射至这些语义描述。推理阶段,通过投影头检索的描述嵌入作为条件,驱动预训练的潜在扩散模型生成图像。该文本中介框架在EEGCVPR数据集上实现了最先进的视觉解码性能,并与已知神经认知通路保持可解释的一致性。主导的EEG-描述关联反映了不同语义层级在感知图像中的重要性。显著性图和t-SNE投影揭示了头皮上的语义拓扑结构。模型展示了结构化语义中介如何实现与认知一致的脑电视觉解码。

原文摘要 · Abstract (English)

Decoding visual experience from brain signals offers exciting possibilities for neuroscience and interpretable AI. While EEG is accessible and temporally precise, its limitations in spatial detail hinder image reconstruction. Our model bypasses direct EEG-to-image generation by aligning EEG signals with multilevel semantic captions -- ranging from object-level to abstract themes -- generated by a large language model. A transformer-based EEG encoder maps brain activity to these captions through contrastive learning. During inference, caption embeddings retrieved via projection heads condition a pretrained latent diffusion model for image generation. This text-mediated framework yields state-of-the-art visual decoding on the EEGCVPR dataset, with interpretable alignment to known neurocognitive pathways. Dominant EEG-caption associations reflected the importance of different semantic levels extracted from perceived images. Saliency maps and t-SNE projections reveal semantic topography across the scalp. Our model demonstrates how structured semantic mediation enables cognitively aligned visual decoding from EEG.

脑机接口图像生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。