用脑区层次结构提升脑影像转图像的精度与可解释性
Hi-DREAM: Brain-Inspired Hierarchical Diffusion for fMRI-to-Image Reconstruction via ROI Encoder and VisuAl Mapping
- 按视觉皮层早期、中期、晚期区域分层编码fMRI信号
- 在自然场景数据集上实现顶尖语义重建效果,保留细节结构
- 可追踪各脑区贡献,适合研究神经机制与生成模型结合
从功能性磁共振成像(fMRI)重建自然图像,需弥合神经活动与现代生成模型所用的结构与语义表征之间的鸿沟。现有基于扩散模型的解码器通常仅依赖单一全局fMRI嵌入,难以利用视觉皮层的层次组织,且各视觉区域的贡献难以分析。我们提出Hi-DREAM,一种受大脑启发的分层扩散框架,依据早期、中期和晚期视觉脑区(ROI)流对fMRI进行分层条件建模。一个ROI适配器将这些流转换为多尺度皮层金字塔,并通过轻量级的ROI条件控制网络,在去噪过程中将解剖感知先验注入匹配的U-Net层级。在自然场景数据集(NSD)上的实验表明,Hi-DREAM在高阶语义重建方面达到当前最优水平,同时保持强低阶结构还原能力。进一步的消融与归因分析显示,所提出的层次感知条件机制有效,不同ROI流提供互补且可解析的重建贡献。
原文摘要 · Abstract (English)
Reconstructing natural images from fMRI requires bridging neural activity with both the structural and semantic representations used by modern generative models. Existing diffusion-based decoders often condition on a single global fMRI embedding, which limits their ability to exploit the hierarchical organization of the visual cortex and makes the contribution of different visual areas difficult to inspect. We propose Hi-DREAM, a brain-inspired hierarchical diffusion framework that structures fMRI conditioning according to early, middle, and late visual Regions of Interest (ROI) streams. A ROI adapter converts these streams into a multi-scale cortical pyramid, and a lightweight ROI-conditioned ControlNet injects the resulting anatomy-aware priors into matched U-Net depths during denoising. Experiments on the Natural Scenes Dataset (NSD) show that Hi-DREAM achieves state-of-the-art high-level semantic reconstruction while retaining strong low-level structure. Further ablation and attribution analyses show that the proposed hierarchy-aware conditioning is effective, and that different ROI streams provide complementary, inspectable contributions to reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。