arXiv:2605.17198q-bio.NCcs.CV2026-05

用视觉数据训练脑活动转图像模型,实现心理图像重建新突破

MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery

论文配图:MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery
图 1 · 摘自论文原文
  • 用线性主干+多模态图文特征驱动扩散模型,专攻心理图像解码
  • 在NSD-Imagery上达成当前最优效果,重构图像与真实高度匹配
  • 适合脑机接口、认知研究领域,为心理成像提供可扩展方案

为支持下游应用,基于脑活动重建所见图像的视觉解码模型必须能泛化到内部生成的视觉表征,即心理图像。对最新发布的NSD-Imagery数据集分析显示,尽管部分现代视觉解码器在心理图像重建上表现良好,但也有模型失败,且在外部刺激图像重建上达到最先进(SOTA)性能,并不保证在心理图像重建上同样领先。受此启发,我们提出MIRAGE,一种专为利用视觉数据集训练并跨模态解码脑活动生成心理图像而设计的方法。MIRAGE采用线性主干结构,以文本和图像多模态特征作为扩散模型输入。通过特征度量和人工评估,MIRAGE在NSD-Imagery基准上确立了当前最优性能。消融分析表明,当解码器使用低维图像特征,并结合文本及高低层次图像特征引导时,心理图像重建效果最佳。本工作表明,只要架构合适,现有大规模外部刺激数据集即可用于心理图像解码训练,预示该领域未来成功与实用性的前景。

原文摘要 · Abstract (English)

To be useful for downstream applications, vision decoding models that are trained to reconstruct seen images from human brain activity must be able to generalize to internally generated visual representations, i.e., mental images. In an analysis of the recently released NSD-Imagery dataset, we demonstrated that while some modern vision decoders can perform quite well on mental image reconstruction, some fail, and that state-of-the-art (SOTA) performance on seen image reconstruction is no guarantee of SOTA performance on mental image reconstruction. Motivated by these findings, we developed MIRAGE, a method explicitly designed to train on vision datasets and cross-decode mental images from brain activity. MIRAGE employs a linear backbone and multi-modal text and image features as input to a diffusion model. Feature metrics and human raters establish MIRAGE as SOTA for mental image reconstruction on the NSD-Imagery benchmark. With ablation analysis we show that mental image reconstruction works best when decoders use image features with relatively few dimensions and include guidance from text-based and both high- and low-level image-based features. Our work indicates that--given the right architecture--existing large-scale datasets using external stimuli are viable training data for decoding mental images, and warrant optimism about the future success and utility of mental image reconstruction.

脑机接口心理图像扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。