语言经验让神经网络的视觉分类表征更分散、语义化。
Language learning shapes visual category-selectivity in deep neural networks
- 用fMRI方法在深度网络中识别出面孔、身体、场景、文字等选择性神经元。
- 语言训练后的模型选择性神经元更多但特异性更低,激活更弱且分布更广。
- 结果暗示语言影响视觉认知,适合关注跨模态学习与脑启发模型的研究者。
人类大脑中的类别选择性区域(如枕叶面孔区FFA、外侧躯体区EBA、海马旁场景区PPA和视觉词形区VWFA)支持高层视觉识别。本文研究人工神经网络是否具备类似的选择性神经元,以及语言经验如何塑造这些表征。通过仿fMRI的功能定位方法,我们在接受类别图像与打乱对照刺激的深度网络中识别出面孔、身体、场景和文字选择性神经元。纯视觉的ResNet与语言监督的Lang-Learned ResNet均在深层出现选择性神经元比例上升的现象。但相比仅视觉模型,语言训练模型表现出更多选择性神经元,但特异性降低,空间定位性减弱,激活强度下降,体现为更分布式、语义对齐的编码模式。这一现象在大规模视觉-语言模型CLIP中得到复现。结果表明,语言经验系统性地重构了神经网络中的视觉类别表征,为语言背景如何影响人脑分类组织提供了计算类比。
原文摘要 · Abstract (English)
Category-selective regions in the human brain-such as the fusiform face area (FFA), extrastriate body area (EBA), parahippocampal place area (PPA), and visual word form area (VWFA)-support high-level visual recognition. Here, we investigate whether artificial neural networks (ANNs) exhibit analogous category-selective neurons and how these representations are shaped by language experience. Using an fMRI-inspired functional localizer approach, we identified face-, body-, place-, and word-selective neurons in deep networks presented with category images and scrambled controls. Both the purely visual ResNet and a linguistically supervised Lang-Learned ResNet contained category-selective neurons that increased in proportion across layers. However, compared to the vision-only model, the Lang-Learned ResNet showed a greater number but lower specificity of category-selective neurons, along with reduced spatial localization and attenuated activation strength-indicating a shift toward more distributed, semantically aligned coding. These effects were replicated in the large-scale vision-language model CLIP. Together, our findings reveal that language experience systematically reorganizes visual category representations in ANNs, providing a computational parallel to how linguistic context may shape categorical organization in the human brain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。