用空间属性训练的视觉模型也能逼近大脑视觉皮层表征。
Vision CNNs trained to estimate spatial latents learned similar ventral-stream-aligned representations
- 用3D图形生成数据,让CNN学习估计物体位置等空间属性。
- 仅学空间属性的模型神经对齐度媲美学数百类别的模型。
- 空间与分类训练模型在早期中层表征高度相似,适合脑科学对照研究。
传统研究认为灵长类背侧视觉通路主要负责物体识别,却忽视了其在估计物体位置、姿态等空间潜变量中的作用。主流背侧通路模型通常通过物体分类任务训练,暗示该通路可能以分类为目标。本文提出新假设:背侧通路是否以估计空间潜变量为优化目标?我们利用3D图形引擎生成的合成图像数据集,训练卷积神经网络(CNN)以估计不同组合的空间和类别潜变量。结果发现,仅训练少数空间潜变量的模型,其神经对齐得分可媲美训练数百类别模型;且空间潜变量表现与神经对齐度强相关。空间与类别训练模型在早期和中间层具有非常相似但不完全相同的内部表征,这种趋同部分由训练数据中非目标潜变量的变异性驱动,促进了对这些未显式训练潜变量的隐式学习。综上,多种训练目标(如空间潜变量)均可产生与背侧通路神经表征对齐的模型。因此,不应默认背侧通路仅优化于物体分类。未来需进一步完善模型与大脑对比的度量方法,以更准确理解背侧通路的功能角色。
原文摘要 · Abstract (English)
Studies of the functional role of the primate ventral visual stream have traditionally focused on object categorization, often ignoring -- despite much prior evidence -- its role in estimating "spatial" latents such as object position and pose. Most leading ventral stream models are derived by optimizing networks for object categorization, which seems to imply that the ventral stream is also derived under such an objective. Here, we explore an alternative hypothesis: Might the ventral stream be optimized for estimating spatial latents? And a closely related question: How different -- if at all -- are representations learned from spatial latent estimation compared to categorization? To ask these questions, we leveraged synthetic image datasets generated by a 3D graphic engine and trained convolutional neural networks (CNNs) to estimate different combinations of spatial and category latents. We found that models trained to estimate just a few spatial latents achieve neural alignment scores comparable to those trained on hundreds of categories, and the spatial latent performance of models strongly correlates with their neural alignment. Spatial latent and category-trained models have very similar -- but not identical -- internal representations, especially in their early and middle layers. We provide evidence that this convergence is partly driven by non-target latent variability in the training data, which facilitates the implicit learning of representations of those non-target latents. Taken together, these results suggest that many training objectives, such as spatial latents, can lead to similar models aligned neurally with the ventral stream. Thus, one should not assume that the ventral stream is optimized for object categorization only. As a field, we need to continue to sharpen our measures of comparing models to brains to better understand the functional roles of the ventral stream.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。