arXiv:2502.18795cs.CL2025-02ACL被引 13

模型能区分自然语言与不可能语言,显示部分类人学习偏好。

Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs

  • 用12种语言测试模型对不可能语法的识别能力
  • 小规模GPT-2可区分自然与不可能语言,但不完美
  • 模型在泛化任务中表现优于困惑度指标

语言模型是否揭示人类语言学习的机制?有人认为,由于架构和训练方式与人类差异巨大,语言模型能轻易学习任意输入,包括非自然语言。我们通过训练模型学习不可能语言和类型学上未见的语言来检验这一观点。不同于以往仅研究英语的研究,本工作在4个语系的12种语言上进行实验,并构建了两个新平行语料库。结果表明,尽管GPT-2 small能够大致区分自然语言与其不可能变体,但无法完全区分开所有自然语言与不可能语言。我们进一步基于格林伯格普遍性20(Greenberg's Universal 20)操纵名词短语顺序,测试模型能否区分类型学上存在与不存在的语序。发现模型的困惑度分数无法区分这些语序,但在泛化测试中表现不同。这表明语言模型具备某些类人的归纳偏置,但其强度弱于人类学习者。

原文摘要 · Abstract (English)

Do language models (LMs) offer insights into human language learning? A common argument against this idea is that because their architecture and training paradigm are so vastly different from humans, LMs can learn arbitrary inputs as easily as natural languages. We test this claim by training LMs to model impossible and typologically unattested languages. Unlike previous work, which has focused exclusively on English, we conduct experiments on 12 languages from 4 language families with two newly constructed parallel corpora. Our results show that while GPT-2 small can largely distinguish attested languages from their impossible counterparts, it does not achieve perfect separation between all the attested languages and all the impossible ones. We further test whether GPT-2 small distinguishes typologically attested from unattested languages with different NP orders by manipulating word order based on Greenberg's Universal 20. We find that the model's perplexity scores do not distinguish attested vs. unattested word orders, while its performance on the generalization test does. These findings suggest that LMs exhibit some human-like inductive biases, though these biases are weaker than those found in human learners.

语言模型认知模拟语法学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。