arXiv:2608.22452cs.CL2026-08

揭示词频与可预测性在儿童语言习得中的不同作用

From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish

论文配图:From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish
图 1 · 摘自论文原文
  • 用词频和可预测性分析西班牙语词汇习得时间
  • 词频显著影响早期词汇掌握,可预测性影响阅读时的处理难度
  • 适合研究语言发展或语言模型评估的学者

surprisal(即语言模型对给定上下文下单词的负对数概率)能可靠预测成人阅读时间。它是否也影响儿童对单个词汇的习得时间?词频反映学习者对词汇的累积暴露,而 surprisal 反映单次出现时的可预测程度。我们通过两项基于语料库的西班牙语研究检验了这一问题。研究1中,使用225个西班牙语名词的词汇频率、儿童语境多样性及三种语言模型(BETO、BERTIN、mGPT)的 surprisal 值,建模其年龄获取值(AoA)。词频强预测 AoA(r = -0.597, p < .001),surprisal 在控制词频和词长后贡献甚微,即使在自然语境分析中亦如此。研究2中,利用 mGPT surprisal 与两个独立频率指标,建模智利西班牙语子样本在 Multilingual Eye-movement Corpus(MECO Wave 2)中的成年读者注视时长。surprisal 在控制频率与词长后仍显著预测更长注视时长,且在两种频率来源下一致。词型层面的匹配比较显示,surprisal 与行为的关联在阅读中强于习得(z = 3.63, p < .001)。结果表明,累积暴露与语境可预测性在语言发展轨迹中扮演不同角色:词频有助于解释早期词汇表征何时形成,而 surprisal 则捕捉已建立语言系统中的即时处理难度。该发现支持基于使用的语言发展理论与扎根理论,并对语言模型作为人类语言行为模型的评估具有启示。

原文摘要 · Abstract (English)

Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words? Frequency reflects a learner's cumulative exposure to a word, whereas surprisal reflects how predictable a single occurrence is given its context. We investigate this question across two corpus-based studies of Spanish. In Study 1, we modeled age of acquisition (AoA) for 225 Spanish nouns using lexical frequency and contextual diversity from child-directed speech, plus surprisal from three language models differing in architecture and training language (BETO, BERTIN, mGPT). Frequency strongly predicted AoA (r=-.597, p<.001); surprisal added little beyond frequency and word length, including in a naturalistic-context analysis. In Study 2, we modeled adult fixation durations in the Chilean Spanish subsample of the Multilingual Eye-movement Corpus (MECO Wave 2), using mGPT surprisal alongside two independent frequency measures. Surprisal robustly predicted longer fixation durations after controlling for frequency and word length, consistent across both frequency sources. A matched word-type-level comparison showed the surprisal-behavior association was stronger in reading than in acquisition (z=3.63, p<.001). The findings suggest cumulative lexical exposure and contextual predictability play different roles across the language trajectory: frequency is particularly informative about when early lexical representations are acquired, whereas surprisal captures moment-to-moment processing difficulty in an already-established linguistic system. We discuss this pattern in relation to usage-based and entrenchment-based accounts of lexical development and to the evaluation of language models as models of human language behavior.

语言发展可预测性词频阅读理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。