用激活最大化生成能精准刺激大脑视觉区的图像。
Visualizing and Controlling Cortical Responses Using Voxel-Weighted Activation Maximization
- 将激活最大化技术用于大脑编码模型,生成可预测脑响应的图像。
- 生成图像能可靠激活目标脑区,涵盖低级与高级视觉皮层。
- 无需生成模型即可灵活调控大脑反应,适合神经科学与脑机接口研究。
在视觉任务上训练的深度神经网络(DNN)发展出类似人脑视觉系统的特征表示。尽管基于DNN的编码模型能准确预测大脑对视觉刺激的响应,但难以揭示驱动这些响应的具体特征。本文展示,原本用于解释视觉DNN的激活最大化技术可应用于基于DNN的大脑编码模型。我们从预训练的Inception V3网络多个层级提取并自适应下采样激活,再通过线性回归预测fMRI响应,构建出全图像可计算的大脑响应模型。随后,对单个皮层体素应用激活最大化,生成优化后能引发预期响应的图像。结果发现,这些图像包含与已知选择性相符的视觉特征,有助于探索整个视觉皮层的选择性。进一步扩展至脑区感兴趣区域(ROIs),并通过人类被试的fMRI实验验证其有效性。结果显示,生成图像能可靠激发目标区域活动,覆盖低、高阶视觉区域及不同受试者。表明激活最大化可成功应用于基于DNN的编码模型,克服了依赖原生生成模型的局限,实现了对人类视觉系统响应的灵活表征与调控。
原文摘要 · Abstract (English)
Deep neural networks (DNNs) trained on visual tasks develop feature representations that resemble those in the human visual system. Although DNN-based encoding models can accurately predict brain responses to visual stimuli, they offer limited insight into the specific features driving these responses. Here, we demonstrate that activation maximization -- a technique designed to interpret vision DNNs -- can be applied to DNN-based encoding models of the human brain. We extract and adaptively downsample activations from multiple layers of a pretrained Inception V3 network, then use linear regression to predict fMRI responses. This yields a full image-computable model of brain responses. Next, we apply activation maximization to generate images optimized for predicted responses in individual cortical voxels. We find that these images contain visual characteristics that qualitatively correspond with known selectivity and enable exploration of selectivity across the visual cortex. We further extend our method to whole regions of interest (ROIs) of the brain and validate its efficacy by presenting these images to human participants in an fMRI study. We find that the generated images reliably drive activity in targeted regions across both low- and high-level visual areas and across subjects. These results demonstrate that activation maximization can be successfully applied to DNN-based encoding models. By addressing key limitations of alternative approaches that require natively generative models, our approach enables flexible characterization and modulation of responses across the human visual system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。