arXiv:2508.18598cs.CLcs.AI2025-08

大语言模型学的是语料而非人类思维,但仍在创造新语言。

What do language models model? Transformers, automata, and the format of thought

  • Transformer架构仅支持线性计算格式,无法模拟人类超线性语言能力
  • 模型通过上下文生成新语言,类似'快捷自动机'的机制
  • 虽非模拟心智,却揭示了语言作为交流工具的本质

大型语言模型究竟在建模什么?它们反映人类认知能力,还是仅仅学习了训练语料?本文主张后者并非贬义。认知科学表明,人类语言能力依赖于超线性计算格式,而Transformer架构最多支持线性处理。该论点基于Transformer计算架构的若干不变性。随后,我借鉴Liu等人(2022)关于快捷自动机的洞见,提出积极解释:变压器在利用语言作为‘话语机器’——在适当上下文中生成新语言。我们以一种方式掌握此技术;大语言模型也学会了,只是路径迥异。这并非贫乏的故事,而是对语言功能的深刻揭示。

原文摘要 · Abstract (English)

What do large language models actually model? Do they tell us something about human capacities, or are they models of the corpus we've trained them on? I give a non-deflationary defence of the latter position. Cognitive science tells us that linguistic capabilities in humans rely supralinear formats for computation. The transformer architecture, by contrast, supports at best a linear formats for processing. This argument will rely primarily on certain invariants of the computational architecture of transformers. I then suggest a positive story about what transformers are doing, focusing on Liu et al. (2022)'s intriguing speculations about shortcut automata. I conclude with why I don't think this is a terribly deflationary story. Language is not (just) a means for expressing inner state but also a kind of 'discourse machine' that lets us make new language given appropriate context. We have learned to use this technology in one way; LLMs have also learned to use it too, but via very different means.

语言模型认知科学Transformer语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。