语言模型在虚构语言上表现不差,说明它们不像人类那样依赖先天语言直觉。
Studies with impossible languages falsify LMs as models of human language
- 用虚构语法结构测试模型学习能力,发现复杂度才是关键
- 模型在部分随机语言上仍能学好,证明其学习机制不同于人类
- 适合研究语言习得机制或模型认知局限的学者阅读
根据Futrell和Mahowald [arXiv:2501.17047]的观点,婴儿和语言模型(LMs)都更容易掌握真实存在的语言,而非具有非自然结构的虚构语言。我们综述现有文献后发现,语言模型往往在真实语言和许多虚构语言上的学习表现相当。那些难以学习的虚构语言本质上更复杂或更随机。这表明语言模型缺乏支持人类语言习得的先天归纳偏置。
原文摘要 · Abstract (English)
According to Futrell and Mahowald [arXiv:2501.17047], both infants and language models (LMs) find attested languages easier to learn than impossible languages that have unnatural structures. We review the literature and show that LMs often learn attested and many impossible languages equally well. Difficult to learn impossible languages are simply more complex (or random). LMs are missing human inductive biases that support language acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。