arXiv:2501.05205cs.CVcs.AI2025-01CVPR被引 2

用婴儿学习数据训练模型,发现它能识别未学过的物体。

Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning

  • 基于婴儿视听数据训练模型,挖掘内部隐藏的视觉概念神经元。
  • 模型可识别超出原始词汇表的物体,展现超越语言的视觉理解能力。
  • 适合对认知科学与视觉模型交叉研究感兴趣的学者。

婴儿在掌握语言前便快速发展出复杂的视觉理解能力。为模拟人类视觉系统,计算机视觉领域可借鉴婴儿视觉发展的机制。本文通过跨学科研究探讨:能否构建一个模仿婴儿学习过程的计算模型,使其发展出超越所听词汇的更广泛视觉概念?我们分析了Vong等人发表于《Science》的模型,该模型基于单个儿童的纵向、第一人称视角图像及父母口语转录文本进行训练。通过对模型内部表示进行神经元标注,我们识别出隐藏的视觉概念神经元,并证明这些神经元能够识别模型未接触过的物体。此外,我们对比了婴儿模型与现代计算机视觉模型(如CLIP和ImageNet预训练模型)在表征上的差异。本研究通过分析基于婴儿视听输入训练的计算模型内部表示,连接认知科学与计算机视觉。

原文摘要 · Abstract (English)

Infants develop complex visual understanding rapidly, even preceding the acquisition of linguistic skills. As computer vision seeks to replicate the human vision system, understanding infant visual development may offer valuable insights. In this paper, we present an interdisciplinary study exploring this question: can a computational model that imitates the infant learning process develop broader visual concepts that extend beyond the vocabulary it has heard, similar to how infants naturally learn? To investigate this, we analyze a recently published model in Science by Vong et al., which is trained on longitudinal, egocentric images of a single child paired with transcribed parental speech. We perform neuron labeling to identify visual concept neurons hidden in the model's internal representations. We then demonstrate that these neurons can recognize objects beyond the model's original vocabulary. Furthermore, we compare the differences in representation between infant models and those in modern computer vision models, such as CLIP and ImageNet pre-trained model. Ultimately, our work bridges cognitive science and computer vision by analyzing the internal representations of a computational model trained on an infant visual and linguistic inputs. Project page is available at https://kexueyi.github.io/webpage-discover-hidden-visual-concepts.

视觉理解婴儿学习神经元分析跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。