arXiv:2507.11522cs.CVcs.LG2025-07中稿 · MICCAI 2025被引 1

用脑电波生成与视觉刺激匹配的图像,准确率超当前最佳13%以上

CATVis: Context-Aware Thought Visualization

  • 五阶段流程:脑电编码+跨模态对齐+重排序+语义融合+图像生成
  • 生成图像与真实视觉刺激对齐度提升,准确率高13.43%,图像质量优36.61%
  • 适合脑机接口、认知解码研究者,尤其关注视觉思维可视化场景

基于脑电的脑机接口在运动想象和认知状态监测中展现出潜力,但从脑电信号解码视觉表征仍面临巨大挑战,因其复杂且噪声大。为此,我们提出一种五阶段新框架:(1) 脑电编码器用于概念分类,(2) 在CLIP特征空间中对齐脑电与文本嵌入,(3) 通过重排序优化字幕,(4) 加权插值概念与字幕嵌入以增强语义,(5) 使用预训练的Stable Diffusion模型生成图像。通过跨模态对齐与重排序实现上下文感知的脑电到图像生成。实验表明,该方法生成的图像与视觉刺激高度一致,分类准确率优于当前最优方法13.43%,生成准确率提升15.21%,弗雷谢特起始距离降低36.61%,表明语义对齐与图像质量显著更优。

原文摘要 · Abstract (English)

EEG-based brain-computer interfaces (BCIs) have shown promise in various applications, such as motor imagery and cognitive state monitoring. However, decoding visual representations from EEG signals remains a significant challenge due to their complex and noisy nature. We thus propose a novel 5-stage framework for decoding visual representations from EEG signals: (1) an EEG encoder for concept classification, (2) cross-modal alignment of EEG and text embeddings in CLIP feature space, (3) caption refinement via re-ranking, (4) weighted interpolation of concept and caption embeddings for richer semantics, and (5) image generation using a pre-trained Stable Diffusion model. We enable context-aware EEG-to-image generation through cross-modal alignment and re-ranking. Experimental results demonstrate that our method generates high-quality images aligned with visual stimuli, outperforming SOTA approaches by 13.43% in Classification Accuracy, 15.21% in Generation Accuracy and reducing Fréchet Inception Distance by 36.61%, indicating superior semantic alignment and image quality.

脑机接口视觉生成跨模态脑电

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。