arXiv:2509.23566cs.CV2025-09中稿 · ICLR被引 1

直接用脑活动信号生成图像,让解码过程更透明可解释。

Towards Interpretable Visual Decoding with Attention to Brain Representations

  • 跳过中间特征空间,直接用脑信号控制扩散模型生成图像。
  • 在公开fMRI数据集上重建质量媲美已有方法。
  • 提出双向可解释框架,揭示不同脑区如何影响生成过程。

近期研究证明,可通过深度生成模型从人类脑活动信号中解码复杂视觉刺激,为探究大脑如何表征真实场景提供了新途径。然而,现有方法通常先将脑信号映射到中间图像或文本特征空间,再引导生成过程,导致各脑区对最终重建的贡献不清晰。本文提出NeuroAdapter,一种直接以脑表征为条件的视觉解码框架,绕过中间特征空间。该方法在公开fMRI数据集上实现了与已有工作相当的重建质量,同时提升了生成过程的透明度。为此,我们引入图像-脑双向可解释性框架(IBBI),分析扩散去噪过程中跨注意力模式,揭示不同皮层区域如何影响生成轨迹。本工作展示了端到端脑到图像重建的潜力,并为可解释神经解码开辟了路径。

原文摘要 · Abstract (English)

Recent work has demonstrated that complex visual stimuli can be decoded from human brain activity using deep generative models, offering new ways to probe how the brain represents real-world scenes. However, many existing approaches first map brain signals into intermediate image or text feature spaces before guiding the generative process, which obscures the contributions of different brain areas to the final reconstruction output. In this work, we propose NeuroAdapter, a visual decoding framework that directly conditions a latent diffusion model on brain representations, bypassing the need for intermediate feature spaces. Our method demonstrates competitive visual reconstruction quality on public fMRI datasets compared to prior work, while providing greater transparency into how brain signals drive visual reconstruction. To this end, we introduce an Image-Brain BI-directional interpretability framework (IBBI) that analyzes cross-attention patterns across diffusion denoising steps to reveal how different cortical areas influence the unfolding generative trajectory. Our work highlights the potential of end-to-end brain-to-image reconstruction and establishes a path for interpretable neural decoding.

脑机接口图像生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。