arXiv:2503.03361cs.AIcs.NE2025-03

模仿婴儿视觉学习机制,提升AI模型的泛化与效率。

Concepts Learned Visually by Infants Can Contribute to Visual Learning and Understanding in AI Models

  • 借鉴婴儿早期习得的动因与目标概念,引导新概念学习。
  • 使用早期概念使模型准确率更高,且仅需更少训练数据。
  • 适合希望提升小样本学习能力的AI研究者参考。

婴幼儿在发育早期能从少量示例中无监督地习得复杂的视觉场景特征,并理解其含义、因果关系及未来事件预测。这些早期概念常用于学习更复杂的概念。本文建模了此类概念如何促进后续学习,并与传统深度网络对比。特别关注动因性(animacy)和目标归属(goal attribution)在动态视觉场景中预测未来事件的作用。结果表明,利用早期概念可显著提升模型准确率并降低数据需求,同时改善表示学习与泛化能力。将先进视觉-语言模型与人类实验对比,在区分有生命与无生命主体行为的任务中,支持早期概念对视觉理解的关键贡献。最后简要讨论将类人视觉学习融入计算机视觉的潜力。

原文摘要 · Abstract (English)

Early in development, infants learn to extract surprisingly complex aspects of visual scenes. This early learning comes together with an initial understanding of the extracted concepts, such as their implications, causality, and using them to predict likely future events. In many cases, this learning is obtained with little or no supervision, and from relatively few examples, compared to current network models. Empirical studies of visual perception in early development have shown that in the domain of objects and human-object interactions, early-acquired concepts are often used in the process of learning additional, more complex concepts. In the current work, we model how early-acquired concepts are used in the learning of subsequent concepts, and compare the results with standard deep network modeling. We focused in particular on the use of the concepts of animacy and goal attribution in learning to predict future events in dynamic visual scenes. We show that the use of early concepts in the learning of new concepts leads to better learning (higher accuracy) and more efficient learning (requiring less data), and that the combination of early and new concepts shapes the representation of the concepts acquired by the model and improves its generalization. We further compare advanced vision-language models to a human study in a task that requires an understanding of the behavior of animate vs. inanimate agents, with results supporting the contribution of early concepts to visual understanding. We finally briefly discuss the possible benefits of incorporating aspects of human-like visual learning into computer vision models.

视觉理解小样本学习类人智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。