arXiv:2603.13696cs.CLcs.LG2026-03

孩子语言模型更爱重复旧词,而非专一对应新词。

Repetition Without Exclusivity: Scale Sensitivity of Referential Mechanisms in Child-Scale Language Models

  • 用重复性提示测试语言模型的指代追踪能力
  • 所有模型均表现出强烈重复倾向,且随训练增强减弱但不消失
  • 结果表明孩子语料难催生词汇排他性,需具身认知支持

我们首次系统评估了仅在儿童话语数据上训练的语言模型中互斥性(ME)的表现——即新词应对应新对象的倾向。将ME操作化为指代抑制:当熟悉物体在双对象语境中被重命名时,预期后续补全中该名词概率下降。三个预实验发现:(1) 掩码模型(BabyBERTa)完全忽略多句指代上下文;(2) 自回归模型对已知名词重命名后表现出显著重复启动(与ME相反);(3) 新诊断工具显示,伪词的类似ME现象实由嵌入相似性解释,非指代消歧。在注册的尺度敏感性实验中,我们训练了45个GPT-2架构模型(290万、890万、3350万参数;在AO-CHILDES上分别训练5、10、20轮,每组5种子),在预注册的ME测试集上评估。所有9组条件下均显著出现反向互斥(重复启动),覆盖85%-100%样本(所有p < 2.4×10^-13)。重复启动随语言建模性能提升而减弱(斯皮尔曼相关rho = -0.533, p = 0.0002),但在3.8倍困惑度范围内始终未归零。上下文依赖诊断在所有9组中复现,且重复次数增加时启动效应增强(8/9组趋势显著,所有p < 0.002)。结果表明,儿童语料上的分布学习产生基于重复的指代追踪,而非词汇排他性。我们将其与具身认知研究联系,认为指代锚定可能是实现互斥性的必要条件,这是一项关于输入结构的经验主张,而非先天论观点。

原文摘要 · Abstract (English)

We present the first systematic evaluation of mutual exclusivity (ME) -- the bias to map novel words to novel referents -- in text-only language models trained on child-directed speech. We operationalise ME as referential suppression: when a familiar object is relabelled in a two-referent discourse context, ME predicts decreased probability of the labelled noun at a subsequent completion position. Three pilot findings motivate a pre-registered scale-sensitivity experiment: (1) a masked language model (BabyBERTa) is entirely insensitive to multi-sentence referential context; (2) autoregressive models show robust repetition priming -- the opposite of ME -- when familiar nouns are re-labelled; and (3) a novel context-dependence diagnostic reveals that apparent ME-like patterns with nonce tokens are fully explained by embedding similarity, not referential disambiguation. In the confirmatory experiment, we train 45 GPT-2-architecture models (2.9M, 8.9M, and 33.5M parameters; 5, 10, and 20 epochs on AO-CHILDES; 5 seeds each) and evaluate on a pre-registered ME battery. Anti-ME repetition priming is significant in all 9 cells (85-100% of items; all p < 2.4 x 10^-13). Priming attenuates with improved language modelling (Spearman rho = -0.533, p = 0.0002) but never crosses zero across a 3.8x perplexity range. The context-dependence diagnostic replicates in all 9 cells, and dose-response priming increases with repetitions in 8/9 cells (all trend p < 0.002). These findings indicate that distributional learning on child-directed speech produces repetition-based reference tracking rather than lexical exclusivity. We connect this to the grounded cognition literature and argue that referential grounding may be a necessary ingredient for ME -- an empirical claim about required input structure, not a nativist one.

语言模型指代理解儿童语料认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。