用大脑分区信号指导图像生成,让脑成像重建更准确
Brain-Streams: fMRI-to-Image Reconstruction with Multi-modal Guidance
- 根据大脑感知与语义分区,分别提取视觉和文本引导信号
- 在真实fMRI数据上实现高保真自然图像重建,细节更丰富
- 适合脑科学、生成模型交叉研究者阅读
理解人类如何处理视觉信息是揭示大脑活动机制的关键一步。近年来,这一好奇心推动了从fMRI数据重建对应视觉刺激的任务:给定视觉刺激下的fMRI数据,目标是还原出原始视觉内容。令人惊讶的是,利用如潜在扩散模型(Latent Diffusion Model, LDM)等强大生成模型,已能在视觉数据集上成功重建高分辨率自然图像。尽管重建图像结构保真度高,但常缺乏小物体细节、模糊形状及语义细微差别。因此,仅依赖视觉信息已不足,需引入额外语义知识。为此,我们借鉴现代LDM能有效融合多模态引导(文本、视觉、图像布局)生成结构与语义合理图像的能力。具体而言,受‘双流假说’启发——感知与语义信息在不同脑区处理——我们的框架Brain-Streams将来自这些脑区的fMRI信号映射为相应嵌入表示。即从语义区域提取文本引导,从感知区域提取视觉引导,从而为LDM提供精准的多模态引导。我们在包含自然图像刺激与对应fMRI数据的真实数据集上,对Brain-Streams的重建能力进行了定量与定性验证。
原文摘要 · Abstract (English)
Understanding how humans process visual information is one of the crucial steps for unraveling the underlying mechanism of brain activity. Recently, this curiosity has motivated the fMRI-to-image reconstruction task; given the fMRI data from visual stimuli, it aims to reconstruct the corresponding visual stimuli. Surprisingly, leveraging powerful generative models such as the Latent Diffusion Model (LDM) has shown promising results in reconstructing complex visual stimuli such as high-resolution natural images from vision datasets. Despite the impressive structural fidelity of these reconstructions, they often lack details of small objects, ambiguous shapes, and semantic nuances. Consequently, the incorporation of additional semantic knowledge, beyond mere visuals, becomes imperative. In light of this, we exploit how modern LDMs effectively incorporate multi-modal guidance (text guidance, visual guidance, and image layout) for structurally and semantically plausible image generations. Specifically, inspired by the two-streams hypothesis suggesting that perceptual and semantic information are processed in different brain regions, our framework, Brain-Streams, maps fMRI signals from these brain regions to appropriate embeddings. That is, by extracting textual guidance from semantic information regions and visual guidance from perceptual information regions, Brain-Streams provides accurate multi-modal guidance to LDMs. We validate the reconstruction ability of Brain-Streams both quantitatively and qualitatively on a real fMRI dataset comprising natural image stimuli and fMRI data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。