探究AI模型与人脑视觉相似性的三大驱动因素
Disentangling the Factors of Convergence between Brains and Computer Vision Models
- 系统性地控制模型规模、训练量和图像类型,研究其对脑相似性的影响
- 大模型+人类中心图像训练可达到最高脑相似度,且存在特定训练时序
- 模型学习顺序与人脑发育顺序一致,揭示了认知发展的潜在规律
许多在自然图像上训练的AI模型会发展出与人脑相似的表征,但其背后的驱动因素尚不明确。为分离模型、训练和数据各自的影响,我们训练了一组自监督视觉变换器(DINOv3),系统性地改变这些因素。我们将模型对图像的表征与通过fMRI和MEG记录的人脑表征进行对比,实现空间与时间上的高分辨率分析。使用三种互补指标评估脑-模型相似性:整体表征相似性、拓扑组织结构、时间动态特征。结果表明,模型规模、训练量和图像类型均独立且交互影响上述三项相似性指标。尤其发现,使用最人类中心图像训练的最大DINOv3模型达到最高脑相似度。模型在训练过程中呈现出特定发展轨迹:先与感觉皮层早期表征对齐,经大量训练后才与晚期及前额叶表征对齐。这一发展过程与人脑皮层的结构和功能特性相关:模型最后获得的表征恰好对应大脑中发育扩展最大、厚度最厚、髓鞘化最少、时间尺度最慢的区域。总体而言,这些发现解耦了架构与经验在塑造人工神经网络如何‘像人一样看世界’中的相互作用,为理解人脑如何表征视觉世界提供了有力框架。
原文摘要 · Abstract (English)
Many AI models trained on natural images develop representations that resemble those of the human brain. However, the factors that drive this brain-model similarity remain poorly understood. To disentangle how the model, training and data independently lead a neural network to develop brain-like representations, we trained a family of self-supervised vision transformers (DINOv3) that systematically varied these different factors. We compare their representations of images to those of the human brain recorded with both fMRI and MEG, providing high resolution in spatial and temporal analyses. We assess the brain-model similarity with three complementary metrics focusing on overall representational similarity, topographical organization, and temporal dynamics. We show that all three factors - model size, training amount, and image type - independently and interactively impact each of these brain similarity metrics. In particular, the largest DINOv3 models trained with the most human-centric images reach the highest brain-similarity. This emergence of brain-like representations in AI models follows a specific chronology during training: models first align with the early representations of the sensory cortices, and only align with the late and prefrontal representations of the brain with considerably more training. Finally, this developmental trajectory is indexed by both structural and functional properties of the human cortex: the representations that are acquired last by the models specifically align with the cortical areas with the largest developmental expansion, thickness, least myelination, and slowest timescales. Overall, these findings disentangle the interplay between architecture and experience in shaping how artificial neural networks come to see the world as humans do, thus offering a promising framework to understand how the human brain comes to represent its visual world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。