arXiv:2412.10489cs.CVcs.AI2024-12AAAI被引 49

用脑电波还原视觉图像,融合多模态信息提升还原精度

CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information

  • 通过多模态专家编码器提取脑电信号中的跨模态特征
  • 利用扩散先验将脑电嵌入映射到CLIP空间,实现高保真图像重建
  • 无需微调生成模型,可扩展至更多模态,适合脑机接口研究者

脑电图(EEG)因其非侵入性和高时间分辨率,在解码视觉刺激方面受到广泛关注。然而,现有研究大多仅关注EEG与图像数据对之间的关系,忽略了EEG信号中蕴含的“图像之外”的多模态信息,导致关键信息丢失。为解决此问题,我们提出CognitionCapturer,一个统一框架,充分利用多模态数据表征EEG信号。具体而言,该框架为每种模态训练模态专家编码器,从EEG模态中提取跨模态信息;随后引入扩散先验,将EEG嵌入空间映射至CLIP嵌入空间,并借助预训练生成模型重建视觉刺激。值得注意的是,该框架无需微调生成模型,且可扩展以融入更多模态。大量实验证明,CognitionCapturer在定性和定量上均优于当前最优方法。

原文摘要 · Abstract (English)

Electroencephalogram (EEG) signals have attracted significant attention from researchers due to their non-invasive nature and high temporal sensitivity in decoding visual stimuli. However, most recent studies have focused solely on the relationship between EEG and image data pairs, neglecting the valuable ``beyond-image-modality" information embedded in EEG signals. This results in the loss of critical multimodal information in EEG. To address this limitation, we propose CognitionCapturer, a unified framework that fully leverages multimodal data to represent EEG signals. Specifically, CognitionCapturer trains Modality Expert Encoders for each modality to extract cross-modal information from the EEG modality. Then, it introduces a diffusion prior to map the EEG embedding space to the CLIP embedding space, followed by using a pretrained generative model, the proposed framework can reconstruct visual stimuli with high semantic and structural fidelity. Notably, the framework does not require any fine-tuning of the generative models and can be extended to incorporate more modalities. Through extensive experiments, we demonstrate that CognitionCapturer outperforms state-of-the-art methods both qualitatively and quantitatively. Code: https://github.com/XiaoZhangYES/CognitionCapturer.

脑机接口图像重建多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。