arXiv:2510.07178cs.CL2025-10Conference of the …被引 3

GPT-2学习自然语言和不可能语言的难度无差别,说明它缺乏人类对语言可能性的敏感性。

Biasless Language Models Learn Unnaturally: How LLMs Fail to Distinguish the Possible from the Impossible

  • 用多种扰动生成不可能语言,测试GPT-2的学习曲线
  • 在多数情况下,自然语言与不可能语言的学习难度相当
  • 模型无法系统区分可能与不可能的语言结构,适合研究认知偏见的读者

大型语言模型(LLMs)是否能区分人类可习得与不可习得的语言?这一问题引发了关于模型与人类是否共享先天学习偏好的讨论。以往研究通过对比现有语言数据集与经不同扰动生成的“不可能”数据集上的学习曲线,得出积极结论。本文采用相同方法,在更广泛的语言及不可能扰动上重新检验该结论。结果显示,大多数情况下,GPT-2学习自然语言及其不可能变体的难易程度并无差异。此外,我们以更宽松的标准评估:基于学习曲线衍生指标的跨语言方差,判断GPT-2能否在整体上区分自然语言与不可能语言。综合两种视角表明,GPT-2未能系统区分可能与不可能的语言。

原文摘要 · Abstract (English)

Are large language models (LLMs) sensitive to the distinction between humanly possible and impossible languages? This question was recently used in a broader debate on whether LLMs and humans share the same innate learning biases. Previous work has answered it in the positive by comparing LLM learning curves on existing language datasets and on "impossible" datasets derived from them via various perturbation functions. Using the same methodology, we examine this claim on a wider set of languages and impossible perturbations. We find that in most cases, GPT-2 learns each language and its impossible counterpart equally easily, in contrast to previous findings. We also apply a more lenient condition by testing whether GPT-2 provides any kind of separation between the whole sets of natural vs. impossible languages, based on cross-linguistic variance in metrics derived from the learning curves. Taken together, these perspectives show that GPT-2 provides no systematic separation between the possible and the impossible.

语言模型认知偏见学习曲线不可能语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。