arXiv:2410.12271cs.CLcs.AI2024-10被引 6

指出大模型学不可能语言实验中的关键漏洞,推动更严谨的语法可学习性研究。

Kallini et al. (2024) do not compare impossible languages with constituency-based ones

  • 揭示现有实验中因语言设计缺陷导致的混淆因素。
  • 证明当前结果无法支持大模型具备人类语言先天约束的结论。
  • 建议改进实验设计,以真正检验语言习得的生物约束机制。

语言理论的核心目标是精确刻画“可能的人类语言”这一概念,即找到一种计算设备,能够描述所有且仅能被正常发育儿童习得的语言。近期大语言模型(LLMs)在自然语言处理中的成功,引发了关于它们是否满足此目标的讨论。若要成立,除了学会人类语言外,还必须难以学会“不可能”的人类语言。Kallini 等人(2024,ACL)通过训练 GPT-2 学习多种合成语言来测试该假设,发现模型对某些语言的学习效果更好。他们将这种不对称性视为大模型归纳偏置与人类语言可能性一致的证据,但最关键的对比存在混淆变量,使结论无效。本文解释了该混淆,并提出改进方向,以构建真正能检验语言可学习性限制的实验。

原文摘要 · Abstract (English)

A central goal of linguistic theory is to find a precise characterization of the notion "possible human language", in the form of a computational device that is capable of describing all and only the languages that can be acquired by a typically developing human child. The success of recent large language models (LLMs) in NLP applications arguably raises the possibility that LLMs might be computational devices that meet this goal. This would only be the case if, in addition to succeeding in learning human languages, LLMs struggle to learn "impossible" human languages. Kallini et al. (2024; "Mission: Impossible Language Models", Proc. ACL) conducted experiments aiming to test this by training GPT-2 on a variety of synthetic languages, and found that it learns some more successfully than others. They present these asymmetries as support for the idea that LLMs' inductive biases align with what is regarded as "possible" for human languages, but the most significant comparison has a confound that makes this conclusion unwarranted. In this paper I explain the confound and suggest some ways forward towards constructing a comparison that appropriately tests the underlying issue.

语言学大模型认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。