arXiv:2606.30815cs.CLcs.AI2026-06

探究Transformer模型学不会人类语言的原因

When transformers learn "impossible" languages, what do they learn?

  • 用扰动英语测试模型语法敏感度,发现性能渐进下降
  • 长序列生成时模型表现严重退化,高质量句子大幅减少
  • 揭示生成能力缺陷是语言无法习得的关键原因

近期研究表明,Transformer语言模型对人类语言存在偏好,而对人为构造的“不可能”语言表现出难以习得的倾向。然而,现有研究主要依赖样本效率和测试困惑度差异,缺乏对可解释人类语言不可习得性的语言能力直接评估。本文检验了两种理论假设:语法敏感性不足或生成能力缺陷。使用GPT-2风格模型在扰动版英语上训练,通过BLiMP最小对测试语法敏感度,发现模型性能仅随语言信息局部性呈渐进下降;但在生成任务中,长序列生成失败显著,高质量句子数量大幅减少。结果表明,生成缺陷与传播失败可能是连接语言模型行为与人类无法习得“不可能”语言的核心机制。

原文摘要 · Abstract (English)

Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to be unacquirable by humans. However, this literature has largely based these claims on differences in sample efficiency and test-set perplexity, rather than on direct evaluations of the linguistic capacities that could plausibly explain non-attestation in human languages. We evaluate two theoretically motivated linking hypotheses: impossibility arising from deficiencies in grammatical sensitivity or generative production. Using GPT-2 style models trained on perturbed "impossible" variants of English, we measure sensitivity to grammaticality using BLiMP minimal pairs, finding that model performance exhibits only gradual degradation, mediated by the language's information locality. In contrast, these models exhibited pronounced failures in generation, producing substantially fewer high-quality sentences at longer lengths. Together, these results suggest generative deficiency and transmission failures as a plausible linking hypothesis between language model behaviour and non-attestation of impossible languages.

语言模型生成缺陷语法敏感性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。