用视觉模型解析大脑如何感知自然图像中的语义特征
Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision Transformers
- 基于预训练视觉模型的密集特征,无须额外训练
- 精准定位大脑高阶视觉皮层对不同视觉概念的响应区域
- 适合研究视觉认知机制或脑机接口的科研人员
我们提出BrainSAIL方法,将神经选择性与自然场景中空间分布的语义视觉概念关联起来。该方法利用大规模人工神经网络提供的密集空间特征,克服自然图像中多类别共现的挑战。通过利用预训练视觉模型的语义一致性特征,并引入一种新型去噪过程,实现无需额外训练的清晰、空间密集的嵌入表示。结合全图嵌入与密集视觉特征,采用体素级编码模型,可识别驱动不同高级视觉皮层区域选择性的具体图像子区域。我们在具有已知类别选择性的皮层区域上验证了该方法,能准确定位并解耦多种视觉概念的选择性。进一步展示其对场景属性及明度、饱和度、深度等低层次特征的高阶选择性表征能力。最后,直接比较不同脑编码模型在视觉皮层各感兴趣区的特征选择性。该方法为解析人类大脑中高层次视觉表征的映射与分解提供了新路径。
原文摘要 · Abstract (English)
We introduce BrainSAIL, a method for linking neural selectivity with spatially distributed semantic visual concepts in natural scenes. BrainSAIL leverages recent advances in large-scale artificial neural networks, using them to provide insights into the functional topology of the brain. To overcome the challenge presented by the co-occurrence of multiple categories in natural images, BrainSAIL exploits semantically consistent, dense spatial features from pre-trained vision models, building upon their demonstrated ability to robustly predict neural activity. This method derives clean, spatially dense embeddings without requiring any additional training, and employs a novel denoising process that leverages the semantic consistency of images under random augmentations. By unifying the space of whole-image embeddings and dense visual features and then applying voxel-wise encoding models to these features, we enable the identification of specific subregions of each image which drive selectivity patterns in different areas of the higher visual cortex. This provides a powerful tool for dissecting the neural mechanisms that underlie semantic visual processing for natural images. We validate BrainSAIL on cortical regions with known category selectivity, demonstrating its ability to accurately localize and disentangle selectivity to diverse visual concepts. Next, we demonstrate BrainSAIL's ability to characterize high-level visual selectivity to scene properties and low-level visual features such as depth, luminance, and saturation, providing insights into the encoding of complex visual information. Finally, we use BrainSAIL to directly compare the feature selectivity of different brain encoding models across different regions of interest in visual cortex. Our innovative method paves the way for significant advances in mapping and decomposing high-level visual representations in the human brain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。