用脑电活动重建图像,同时保持语义和结构准确。
Versatile Framework with Semantic and Structural guidance for Image Reconstruction from Brain Activity

- 分两阶段:先用CLIP生成语义图像,再用视觉特征迭代优化结构
- 在三种脑信号数据上均超越现有模型,提升图像一致性
- 适合脑机接口、神经解码研究者,推动可解释性重建
从脑记录中重构视觉刺激是一项意义重大且具有挑战性的任务,尤其精确可控的图像重建对脑机接口的发展至关重要。近期方法借助文本到图像生成模型的进展,在语义层面(如概念和物体)实现了与复杂自然刺激的高接近度。然而,这些方法在细粒度结构信息(如位置、方向和尺寸)上难以保持一致,削弱了模型的可控性与可解释性。为此,我们提出名为MindDiffuser的两阶段图像重建框架。第一阶段,将从脑响应中解码出的对比语言-图像预训练(CLIP)文本嵌入输入Stable Diffusion,生成包含语义信息的初步图像。第二阶段,利用解码的浅层CLIP视觉特征作为监督信号,通过反向传播迭代优化第一阶段的特征向量,以对齐结构信息。我们在三种模态(fMRI、EEG、MEG)的脑响应数据集上进行了广泛实验,结果表明该框架显著提升了先前最先进模型的性能,验证了方法的有效性与通用性。空间与时间可视化结果进一步支持了框架的神经生物学合理性,为跨模态神经解码研究提供了指导。
原文摘要 · Abstract (English)
Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task in brain decoding. Especially, the achievement of precise and controllable image reconstruction bears great significance in propelling the progress and utilization of brain-computer interfaces. Recent methods, leveraging advances in the power of text-to-image generation models, have reconstructed images that closely approximate complex natural stimuli in terms of semantics (e.g., concepts and objects). However, they struggle to maintain consistency with the original stimuli in fine-grained structural information (e.g., position, orientation and size), which undermines both the controllability and interpretability of the models. To address the aforementioned issues, we propose a two-stage image reconstruction framework, termed MindDiffuser. In Stage 1, Contrastive Language-Image Pretraining (CLIP) text embeddings decoded from brain responses are input into Stable Diffusion, generating a preliminary image containing semantic information. In Stage 2, we use decoded shallow CLIP visual features as supervisory signals, iteratively refining the feature vectors from Stage 1 via backpropagation to align structural information. We conducted extensive experiments on brain response datasets across three modalities (fMRI, EEG, MEG) elicited by visual stimuli, demonstrating that our framework significantly enhances the performance of previous state-of-the-art models, highlighting the effectiveness and versatility of our approach. Spatial and temporal visualization results further support the neurobiological plausibility of our framework, providing guidance for future neural decoding efforts across different brain signal modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。