arXiv:2506.02147cs.CL2025-06中稿 · EMNLP被引 1

小模型在合理数据量下仍能学会复杂语言构造,表现与任务成绩相关。

BabyLM's First Constructions: Causal probing provides a signal of learning

  • 用因果探针检测婴儿语言模型对语言构造的学习情况
  • 仅用发展性合理的数据量,模型已掌握多种构造,包括难辨类型
  • 构造表征能力越强,在婴儿语言挑战赛中表现越好

建构语法认为语言学习者从环境统计中习得构式(形式-语义配对)。近期研究支持此观点,发现预训练语言模型对构式敏感,如Rozner等(2025)证明构式影响RoBERTa的输出分布。但这些模型通常在发展上不合理的海量数据上训练,使其与人类语言学习的相关性存疑。本文采用Rozner等的方法,评估2024年婴儿语言模型挑战赛中掩码语言模型的构式学习能力。结果表明,即使在发展性合理的数据量下,模型仍能学习多样构式,包括表面难以区分的难点构式。我们还发现构式表现与任务成绩存在相关性:构式表征能力越强,模型在BabyLM基准测试中表现越好。

原文摘要 · Abstract (English)

Construction grammar posits that language learners acquire constructions (form-meaning pairings) from the statistics of their environment. Recent work supports this hypothesis by showing sensitivity to constructions in pretrained language models (PLMs), including one recent study (Rozner et al., 2025) demonstrating that constructions shape RoBERTa's output distribution. However, models under study have generally been trained on developmentally implausible amounts of data, casting doubt on their relevance to human language learning. Here we use Rozner et al.'s methods to evaluate construction learning in masked language models from the 2024 BabyLM Challenge. Our results show that even when trained on developmentally plausible quantities of data, models learn diverse constructions, even hard cases that are superficially indistinguishable. We further find correlational evidence that constructional performance may be functionally relevant: models that better represent construction perform better on the BabyLM benchmarks.

语言模型构式语法学习机制婴儿语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。