CDL训练未必提升语法学习,大模型在多个语言中表现不如维基数据
Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models
- 对比儿童语料与维基百科训练模型的语法能力
- 多数情况下维基模型优于儿童语料模型,且效果不一致
- 提出新评测方法FIT-CLAMS,控制频率偏差
Huebner等(2021)的开创性研究发现,用英语儿童语料(CDL)训练的语言模型可达到与使用大量成人文本训练模型相当的句法能力,暗示CDL可能比常用互联网爬取数据更有效。然而,这一结论在不同语言、模型类型和评估设置下的普适性尚不明确。本文通过对比在CDL与维基百科上训练的模型,在两种语言模型目标(掩码与自回归)、三种语言(英语、法语、德语)及三个句法最小对差异基准上进行测试。结果显示,CDL在大多数情况下未能带来稳定优势,反而常被维基模型超越。我们进一步指出此前评测方法的多种缺陷,并提出新型评测框架FIT-CLAMS,采用频率控制设计实现训练语料间的公平比较。通过最小对评估与回归分析,证实CDL训练并未增强句法泛化能力,强调在评估句法能力时必须控制频率效应。
原文摘要 · Abstract (English)
Seminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of adult-directed written text, suggesting that CDL could provide more effective LM training material than the commonly used internet-crawled data. However, the generalizability of these results across languages, model types, and evaluation settings remains unclear. We test this by comparing models trained on CDL vs. Wikipedia across two LM objectives (masked and causal), three languages (English, French, German), and three syntactic minimal-pair benchmarks. Our results on these benchmarks show inconsistent benefits of CDL, which in most cases is outperformed by Wikipedia models. We then identify various shortcomings in previous benchmarks, and introduce a novel testing methodology, FIT-CLAMS, which uses a frequency-controlled design to enable balanced comparisons across training corpora. Through minimal pair evaluations and regression analysis we show that training on CDL does not yield stronger generalizations for acquiring syntax and highlight the importance of controlling for frequency effects when evaluating syntactic ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。