arXiv:2505.16419cs.CVcs.AI2025-05被引 6

用新方法对比神经网络与人脑对物体的相似判断,发现CLIP模型更像人脑。

Investigating Fine- and Coarse-grained Structural Correspondences Between Deep Neural Networks and Human Object Image Similarity Judgments Using Unsupervised Alignment

  • 用无监督对齐方法找人脑与模型间物体表示的最优对应关系
  • CLIP模型在细粒度和粗粒度上均与人脑高度匹配,自监督模型仅粗粒度有效
  • 揭示语言信息对精确物体表征的重要性,适合认知科学与模型评估研究者

人类如何形成物体内部表征的学习机制尚不明确。深度神经网络(DNN)因优化目标函数后生成类似人脑的内部表征,成为研究该问题的有力工具。尽管已有研究显示,经不同学习范式(如监督、自监督、CLIP)训练的模型可获得类人表征,但其与人脑的相似性是仅限于粗粒度类别,还是延伸至细粒度细节仍不清楚。本文采用基于Gromov-Wasserstein最优传输的无监督对齐方法,在细粒度和粗粒度层面比较人类与模型的物体表征。该方法相比传统表示相似性分析的独特优势在于能估计每个物体在人类与模型表征间的最优细粒度映射。利用来自THINGS数据集的1,854个物体的人类相似性判断,我们发现:使用CLIP训练的模型在细粒度与粗粒度层面均与人类表征保持强一致性;而自监督模型在两个层次上的匹配程度有限,但仍能形成反映人类粗粒度类别结构的物体聚类。结果为语言信息在构建精确物体表征中的作用提供了新见解,并展示了自监督学习在捕捉粗粒度类别结构方面的潜力。

原文摘要 · Abstract (English)

The learning mechanisms by which humans acquire internal representations of objects are not fully understood. Deep neural networks (DNNs) have emerged as a useful tool for investigating this question, as they have internal representations similar to those of humans as a byproduct of optimizing their objective functions. While previous studies have shown that models trained with various learning paradigms - such as supervised, self-supervised, and CLIP - acquire human-like representations, it remains unclear whether their similarity to human representations is primarily at a coarse category level or extends to finer details. Here, we employ an unsupervised alignment method based on Gromov-Wasserstein Optimal Transport to compare human and model object representations at both fine-grained and coarse-grained levels. The unique feature of this method compared to conventional representational similarity analysis is that it estimates optimal fine-grained mappings between the representation of each object in human and model representations. We used this unsupervised alignment method to assess the extent to which the representation of each object in humans is correctly mapped to the corresponding representation of the same object in models. Using human similarity judgments of 1,854 objects from the THINGS dataset, we find that models trained with CLIP consistently achieve strong fine- and coarse-grained matching with human object representations. In contrast, self-supervised models showed limited matching at both fine- and coarse-grained levels, but still formed object clusters that reflected human coarse category structure. Our results offer new insights into the role of linguistic information in acquiring precise object representations and the potential of self-supervised learning to capture coarse categorical structures.

神经网络认知科学表征学习对比分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。