用结构化文本空间提升脑影像转图像的精度,关键发现是文本表征更贴近大脑视觉活动。
Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- 将fMRI信号映射到结构化文本空间,作为图像重建的中间表示
- 在真实数据集上实现感知损失降低8%,优于现有方法
- 适合对神经编码机制和生成模型融合感兴趣的跨学科研究者
理解大脑如何编码视觉信息是神经科学与机器学习的核心挑战。一种有前景的方法是从功能磁共振成像(fMRI)信号中重构视觉刺激,即图像。该过程包含两个阶段:将fMRI信号转换为潜在空间,再通过预训练生成模型重构图像。重建质量取决于潜在空间与神经活动结构的相似性,以及生成模型从该空间生成图像的能力。然而,何种潜在空间最适配此任务仍不明确。本文提出两大发现:首先,fMRI信号与语言模型的文本空间更相似,而非视觉或图文联合空间;其次,文本表征与生成模型需适应视觉刺激的组合性,包括物体、属性及关系。基于此,我们提出PRISM模型,将fMRI信号投影至结构化文本空间作为中间表示。其包含基于对象中心的扩散模块,通过组合独立物体减少检测错误;以及属性关系搜索模块,自动识别与神经活动最匹配的关键属性与关系。在真实世界数据集上的大量实验表明,本框架显著优于现有方法,感知损失最高降低8%。结果凸显了使用结构化文本作为桥梁连接fMRI信号与图像重建的重要性。
原文摘要 · Abstract (English)
Understanding how the brain encodes visual information is a central challenge in neuroscience and machine learning. A promising approach is to reconstruct visual stimuli, essentially images, from functional Magnetic Resonance Imaging (fMRI) signals. This involves two stages: transforming fMRI signals into a latent space and then using a pretrained generative model to reconstruct images. The reconstruction quality depends on how similar the latent space is to the structure of neural activity and how well the generative model produces images from that space. Yet, it remains unclear which type of latent space best supports this transformation and how it should be organized to represent visual stimuli effectively. We present two key findings. First, fMRI signals are more similar to the text space of a language model than to either a vision based space or a joint text image space. Second, text representations and the generative model should be adapted to capture the compositional nature of visual stimuli, including objects, their detailed attributes, and relationships. Building on these insights, we propose PRISM, a model that Projects fMRI sIgnals into a Structured text space as an interMediate representation for visual stimuli reconstruction. It includes an object centric diffusion module that generates images by composing individual objects to reduce object detection errors, and an attribute relationship search module that automatically identifies key attributes and relationships that best align with the neural activity. Extensive experiments on real world datasets demonstrate that our framework outperforms existing methods, achieving up to an 8% reduction in perceptual loss. These results highlight the importance of using structured text as the intermediate space to bridge fMRI signals and image reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。