arXiv:2507.12009cs.CVcs.HC2025-07中稿 · International Conf…

用深度模型从脑活动重建电影画面,揭示视觉处理关键区域。

Deep Neural Encoder-Decoder Model to Relate fMRI Brain Activity with Naturalistic Stimuli

  • 用时序卷积网络建模连续影像与脑活动的对应关系。
  • 可重建画面中的边缘、人脸和对比度等视觉特征。
  • 发现枕叶中区、梭状回和楔前叶是主要贡献区域。

我们提出一种端到端的深度神经编码-解码模型,利用功能磁共振成像(fMRI)数据,将大脑对自然影像刺激的响应进行编码与解码。通过引入连续电影帧的时间相关性,模型采用时序卷积层,有效弥补了自然电影刺激与fMRI采样之间的时间分辨率差距。该模型能预测视觉皮层及其周边体素的活动,并从神经活动中重构对应的视觉输入。此外,我们通过显著性图分析了参与视觉解码的大脑区域,发现最活跃区域分别为中枕区(形状感知)、梭状回(复杂识别,尤其是面孔识别)和楔前叶(基本视觉特征如边缘与对比度)。这些功能与解码器成功重建边缘、人脸和对比度的能力高度一致。总体表明,可通过此类深度学习模型,以代理方式探究电影中视觉加工的神经机制。

原文摘要 · Abstract (English)

We propose an end-to-end deep neural encoder-decoder model to encode and decode brain activity in response to naturalistic stimuli using functional magnetic resonance imaging (fMRI) data. Leveraging temporally correlated input from consecutive film frames, we employ temporal convolutional layers in our architecture, which effectively allows to bridge the temporal resolution gap between natural movie stimuli and fMRI acquisitions. Our model predicts activity of voxels in and around the visual cortex and performs reconstruction of corresponding visual inputs from neural activity. Finally, we investigate brain regions contributing to visual decoding through saliency maps. We find that the most contributing regions are the middle occipital area, the fusiform area, and the calcarine, respectively employed in shape perception, complex recognition (in particular face perception), and basic visual features such as edges and contrasts. These functions being strongly solicited are in line with the decoder's capability to reconstruct edges, faces, and contrasts. All in all, this suggests the possibility to probe our understanding of visual processing in films using as a proxy the behaviour of deep learning models such as the one proposed in this paper.

脑机接口视觉解码fMRI深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。