arXiv:2505.17329q-bio.NCcs.LG2025-05NeurIPS被引 11

用Transformer注意力机制解释大脑如何动态分配视觉信息

Transformer brain encoders explain human high-level visual responses

  • 引入Transformer注意力机制模拟大脑高阶区域的动态信息路由
  • 在自然场景下预测大脑活动表现优于传统方法,显著提升准确率
  • 结果可直观可视化,适合研究视觉认知机制与脑-机接口

神经科学的核心目标之一是理解大脑在自然场景下进行视觉处理的计算机制。当前主流方法使用可图像计算的深度神经网络,通过线性编码模型拟合脑活动,但需估计大量参数,且忽略脑区与模型特征图的结构关联。近期方法将线性映射分解为空间与特征权重,仅适用于早期视觉皮层。本文采用Transformer架构中的注意力机制,研究视网膜拓扑特征如何被动态路由至高阶类别选择性脑区。实验表明,该方法在多种特征基模型和模态下,均显著优于现有方法,能更准确预测自然场景观看时的大脑反应。此外,注意力路由信号可直接可视化,具有更高可解释性。由于其对新图像的预测性能优异,该模型有望成为视觉信息从视网膜图谱向类别选择区域动态传递的候选机制模型。

原文摘要 · Abstract (English)

A major goal of neuroscience is to understand brain computations during visual processing in naturalistic settings. A dominant approach is to use image-computable deep neural networks trained with different task objectives as a basis for linear encoding models. However, in addition to requiring estimation of a large number of linear encoding parameters, this approach ignores the structure of the feature maps both in the brain and the models. Recently proposed alternatives factor the linear mapping into separate sets of spatial and feature weights, thus finding static receptive fields for units, which is appropriate only for early visual areas. In this work, we employ the attention mechanism used in the transformer architecture to study how retinotopic visual features can be dynamically routed to category-selective areas in high-level visual processing. We show that this computational motif is significantly more powerful than alternative methods in predicting brain activity during natural scene viewing, across different feature basis models and modalities. We also show that this approach is inherently more interpretable as the attention-routing signals for different high-level categorical areas can be easily visualized for any input image. Given its high performance at predicting brain responses to novel images, the model deserves consideration as a candidate mechanistic model of how visual information from retinotopic maps is routed in the human brain based on the relevance of the input content to different category-selective regions.

Transformer脑科学视觉认知注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。