用脑区交互模型从fMRI重建图像,更准且只需少量数据
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
- 设计脑区交互变压器,让相似脑区协同处理信息
- 1小时数据即可达40小时训练效果,重建图像更贴近真实
- 同时预测语义与结构特征,提升生成图像质量
从人观看图像时的fMRI脑信号中重建图像,为非侵入式理解大脑提供窗口。尽管扩散模型带来进展,现有方法常缺乏对真实图像的忠实还原。我们提出Brain-IT,基于脑交互变换器(BIT),实现功能相似脑体素簇间的有效交互。这些体素簇在所有受试者间共享,作为整合脑内与跨脑信息的基本单元。所有模型组件对所有簇和受试者共享,可在有限数据下高效训练。为引导图像重建,BIT预测两类互补的局部图像特征:(i) 高层语义特征,指导扩散模型生成正确语义内容;(ii) 低层结构特征,帮助以正确粗略布局初始化扩散过程。该设计实现脑区簇到图像局部特征的直接信息流。基于此,本方法在视觉效果和标准客观指标上均优于当前最先进方法,且仅需新受试者1小时的fMRI数据,即达到以往需40小时数据训练的效果。
原文摘要 · Abstract (English)
Reconstructing images seen by people from their fMRI brain recordings provides a non-invasive window into the human brain. Despite recent progress enabled by diffusion models, current methods often lack faithfulness to the actual seen images. We present "Brain-IT", a brain-inspired approach that addresses this challenge through a Brain Interaction Transformer (BIT), allowing effective interactions between clusters of functionally-similar brain-voxels. These functional-clusters are shared by all subjects, serving as building blocks for integrating information both within and across brains. All model components are shared by all clusters & subjects, allowing efficient training with a limited amount of data. To guide the image reconstruction, BIT predicts two complementary localized patch-level image features: (i)high-level semantic features which steer the diffusion model toward the correct semantic content of the image; and (ii)low-level structural features which help to initialize the diffusion process with the correct coarse layout of the image. BIT's design enables direct flow of information from brain-voxel clusters to localized image features. Through these principles, our method achieves image reconstructions from fMRI that faithfully reconstruct the seen images, and surpass current SotA approaches both visually and by standard objective metrics. Moreover, with only 1-hour of fMRI data from a new subject, we achieve results comparable to current methods trained on full 40-hour recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。