儿童语调促进大模型语言生成,而非理解能力。
Child-directed speech facilitates production, not comprehension, in BabyLMs
- 设计框架填空任务评估生成能力,契合使用基础语言习得理论。
- 儿童语调训练模型在早期就更准确完成句子框架,概率集中于合理填充词。
- 适用于研究儿童语言学习机制或提升模型生成能力的学者。
近期研究认为儿童语调(CDS)对婴儿语言模型(BabyLMs)的语言学习无益,但现有评估主要关注理解能力,忽视了生成能力——这正是使用基础理论强调的核心。本文提出一种基于生成的任务,模拟语言习得中的结构化‘框架’(频繁出现的词汇模式带开放槽位)。我们对比了在儿童语调数据、BabyLM语料库及网页爬取数据(FineWeb-edu)上训练的Llama模型,在理解基准与新框架填空任务上的表现。结果揭示:尽管以FineWeb训练的模型在最小对辨识任务中表现更优,但以儿童语调训练的模型在训练初期即能更早生成语法正确的句子,且概率更集中于恰当的槽位填充词。表明当前理解类评估低估了儿童语调对婴儿语言模型生成能力的促进作用。
原文摘要 · Abstract (English)
Recent studies suggest that child-directed speech is not conducive to language learning in BabyLMs. However, current evaluations focus predominantly on comprehension and not production, which is central to usage-based theories of language acquisition which argue how CDS facilitates early language use through constructional ''frames'' (frequent lexical patterns with open slots). We introduce a novel generation-based evaluation inspired by such theories in form of a frame-completion task, and compare Llama models trained with CDS, the BabyLM corpus, and web-crawl data (FineWeb-edu) on comprehension benchmarks and our novel framework. Our results reveal a clear dissociation between models' comprehension and production capabilities: while FineWeb-trained models excel at minimal pairs, CDS-trained models produce grammatical completions substantially earlier in training and concentrate probability mass on appropriate slot-fillers. These findings show that comprehension benchmarks underestimate what CDS affords to BabyLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。