arXiv:2603.29522cs.CLcs.AI2026-03被引 2

用婴儿语言数据训练模型,发现真实语料更利于语法学习

Baby Scale: Investigating Models Trained on Individual Children's Language Input

  • 在婴儿自然语言数据上训练模型,考察小规模数据下的学习规律
  • 模型在语法任务上表现尚可,但语义和世界知识任务效果差于合成数据
  • 语言互动特征比数据量更能预测学习效果,适合研究儿童语言发展

现代语言模型需远超儿童实际接收的语料量才能产生有效行为。为理解这一‘数据差距’的本质,我们使用来自6-36个月婴儿的BabyView数据集视频转录文本,研究三方面问题:(1)在儿童尺度数据下模型的缩放性能;(2)不同儿童数据集间的性能差异及语言特征对数据质量的预测能力;(3)模型与儿童语言学习结果的关系。结果显示,基于儿童数据训练的模型在语法任务上表现出可接受的缩放规律,但在语义和世界知识任务上缩放效果低于合成数据训练模型;且不同儿童数据间存在显著性能差异。除数据量外,分布性和交互性语言特征组合最能预测模型表现,与促进儿童语言发展的优质输入一致。此外,模型对单个词的似然值与儿童对该词的学习程度相关,表明亲子语言输入特性可能同时影响模型与人类的语言学习。总体而言,识别高效语言学习的数据属性,有助于构建更强的小规模语言模型,并揭示人类语言习得机制。

原文摘要 · Abstract (English)

Modern language models (LMs) must be trained on many orders of magnitude more words of training data than human children receive before they begin to produce useful behavior. Assessing the nature and origins of this "data gap" requires benchmarking LMs on human-scale datasets to understand how linguistic knowledge emerges from children's natural training data. Using transcripts from the BabyView dataset (videos from children ages 6-36 months), we investigate (1) scaling performance at child-scale data regimes, (2) variability in model performance across datasets from different children's experiences and linguistic predictors of dataset quality, and (3) relationships between model and child language learning outcomes. LMs trained on child data show acceptable scaling for grammar tasks, but lower scaling on semantic and world knowledge tasks than models trained on synthetic data; we also observe substantial variability on data from different children. Beyond dataset size, performance is most associated with a combination of distributional and interactional linguistic features, broadly consistent with what makes high-quality input for child language development. Finally, model likelihoods for individual words correlate with children's learning of those words, suggesting that properties of child-directed input may influence both model learning and human language development. Overall, understanding what properties make language data efficient for learning can enable more powerful small-scale language models while also shedding light on human language acquisition.

语言模型儿童语言数据效率学习机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。