用AI分析婴儿视角视频,发现语言与视觉匹配机会很少。
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
- 用CLIP模型自动分析婴儿视角视频中的视听同步性
- 发现真实生活中语言与物体匹配时刻占比不足10%
- 为儿童语言学习研究提供新方法,适合发展心理学与AI交叉研究者
理解词语指代何物是幼儿语言学习的核心挑战。现有模型认为儿童通过日常环境中说话人提及物体时的词物共现来习得词汇。但儿童的视觉与语言体验在时间上是否对齐?以往研究受限于人工标注成本。本文利用对比语言-图像预训练(CLIP)模型,自动评估家庭环境中婴儿第一视角视频中的视听对齐程度。经人工判断验证后,应用于大规模婴儿视角视频数据集。结果显示,理想化的学习对齐时刻(如说“看球”时球确实在视野中)在真实日常经验中相对罕见,远低于现代机器学习数据集水平,且不同儿童间及同一儿童内部均存在显著差异。这表明低频对齐可能是早期词汇学习模型的重要限制因素,并提出一种研究儿童多模态环境的新方法。
原文摘要 · Abstract (English)
Figuring out which objects or concepts words refer to is a central language learning challenge for young children. Most models of this process posit that children learn early object labels from co-occurrences of words and their referents that occur when someone around them talks about an object in the immediate physical environment. But how aligned in time are children's visual and linguistic experiences during everyday learning? To date, answers to this question have been limited by the need for labor-intensive manual annotations of vision-language co-occurrences. Here, we evaluate the use of contrastive language-image pretraining (CLIP) models to automatically characterize vision-language alignment in egocentric videos taken from the infant perspective in home environments. After validating CLIP alignment scores using human alignment judgments, we apply this metric to a large corpus of infant-perspective videos. We show that idealized aligned moments for learning (e.g., "look at the ball" with a ball present in the child's view) are relatively rare in children's everyday experiences compared to modern machine learning datasets, and highlight variability in alignment both within and across children. These findings suggest that infrequent alignment is a constraint for models describing early word learning and offer a new method for investigating children's multimodal environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。