arXiv:2505.13281cs.CV2025-05被引 1

计算机视觉模型对几何拓扑概念的敏感度接近人类儿童,且表现与人类一致。

Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts

  • 用三类视觉模型测试43个几何拓扑概念的识别能力。
  • 基于Transformer的模型准确率超过儿童,且与人类难度偏好高度一致。
  • 融合语言视觉的模型反而表现更差,提示多模态可能损害抽象感知。

随着机器学习模型的快速发展,认知科学家愈发关注其与人类思维的契合度。本文探讨计算机视觉模型与人类对几何与拓扑(GT)概念的敏感性是否一致。根据核心知识理论,这些概念是先天具备的,由专门神经回路支持。本文提出另一种解释:GT概念可通过日常环境互动“免费”习得。我们基于大规模图像数据集训练的三类模型——卷积神经网络(CNN)、基于Transformer的模型和视觉-语言模型——在包含43个概念、覆盖七类的奇偶识别任务中进行测试。基于Transformer的模型总体准确率最高,超过年幼儿童;且在哪些概念易、哪些难的问题上,与儿童表现高度一致。相比之下,视觉-语言模型的表现低于纯视觉模型,并偏离人类模式更远,表明简单的多模态融合可能损害抽象几何敏感性。这些结果支持用计算机视觉模型检验‘学习假说’的充分性,同时暗示语言与视觉表征的整合可能带来未预期的负面效应。

原文摘要 · Abstract (English)

With the rapid improvement of machine learning (ML) models, cognitive scientists are increasingly asking about their alignment with how humans think. Here, we ask this question for computer vision models and human sensitivity to geometric and topological (GT) concepts. Under the core knowledge account, these concepts are innate and supported by dedicated neural circuitry. In this work, we investigate an alternative explanation, that GT concepts are learned ``for free'' through everyday interaction with the environment. We do so using computer visions models, which are trained on large image datasets. We build on prior studies to investigate the overall performance and human alignment of three classes of models -- convolutional neural networks (CNNs), transformer-based models, and vision-language models -- on an odd-one-out task testing 43 GT concepts spanning seven classes. Transformer-based models achieve the highest overall accuracy, surpassing that of young children. They also show strong alignment with children's performance, finding the same classes of concepts easy vs. difficult. By contrast, vision-language models underperform their vision-only counterparts and deviate further from human profiles, indicating that naïve multimodality might compromise abstract geometric sensitivity. These findings support the use of computer vision models to evaluate the sufficiency of the learning account for explaining human sensitivity to GT concepts, while also suggesting that integrating linguistic and visual representations might have unpredicted deleterious consequences.

计算机视觉几何感知多模态模型人类认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。