arXiv:2507.14749cs.CL2025-07被引 4

用单个儿童的有限数据训练模型,也能让机器学会词与物的对应关系。

On the robustness of modeling grounded word learning through a child's egocentric input

  • 用自动化转录技术处理500小时儿童视频,构建多模态训练数据
  • 三个孩子各自的数据都能让模型学到词与物体的关联,且跨数据泛化
  • 适合研究儿童语言习得机制的学者,也适用于多模态学习应用

机器学习能为理解人类语言习得提供什么洞见?尽管大语言和多模态模型表现优异,但其依赖海量数据的特点与儿童仅凭有限输入就能习得语言的事实存在根本矛盾。为弥合这一差距,研究者开始使用接近儿童实际输入量级的数据训练神经网络。Vong等(2024)曾证明,仅用一个儿童61小时的视听数据训练的多模态网络即可实现词-指称映射学习。但这种成功是否源于个别儿童的独特经验,还是在多个儿童数据上均具鲁棒性,尚不清楚。本文对包含三名儿童、总计超过500小时视频的SAYCam数据集进行全量自动化语音转录,构建用于训练与评估的多模态视觉-语言数据集,并测试多种神经网络配置,以评估模拟词义学习的鲁棒性。结果表明,基于每位儿童自动转录数据训练的模型均可习得词-指称映射,并在跨视频、跨儿童及跨图像领域中实现泛化。这些发现验证了多模态神经网络在具身词学习中的鲁棒性,同时揭示了不同个体发展经验下模型学习路径的差异。

原文摘要 · Abstract (English)

What insights can machine learning bring to understanding human language acquisition? Large language and multimodal models have achieved remarkable capabilities, but their reliance on massive training datasets creates a fundamental mismatch with children, who succeed in acquiring language from comparatively limited input. To help bridge this gap, researchers have increasingly trained neural networks using data similar in quantity and quality to children's input. Taking this approach to the limit, Vong et al. (2024) showed that a multimodal neural network trained on 61 hours of visual and linguistic input extracted from just one child's developmental experience could acquire word-referent mappings. However, whether this approach's success reflects the idiosyncrasies of a single child's experience, or whether it would show consistent and robust learning patterns across multiple children's experiences was not explored. In this article, we applied automated speech transcription methods to the entirety of the SAYCam dataset, consisting of over 500 hours of video data spread across all three children. Using these automated transcriptions, we generated multi-modal vision-and-language datasets for both training and evaluation, and explored a range of neural network configurations to examine the robustness of simulated word learning. Our findings demonstrate that networks trained on automatically transcribed data from each child can acquire word-referent mappings, generalizing across videos, children, and image domains. These results validate the robustness of multimodal neural networks for grounded word learning, while highlighting the individual differences that emerge in how models learn when trained on each child's developmental experiences.

语言习得多模态学习儿童发展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。