arXiv:2606.02953cs.CL2026-06

大模型能模仿语言创造性,但不会因没见过的结构而避免错误使用。

Linguistic Productivity in Large Language Models: Models Coerce, but do not Preempt

  • 通过大量文本学习的语言模型具备构造性生成能力。
  • 大模型能正确使用生造词的非常规语义,体现语言创造力。
  • 模型无法从缺失数据中学习规避错误,仍会过度泛化。

基于使用理论的语法认为,语言的创造性既受高频使用的巩固作用(固有化)推动,也受未在预期场景中出现过结构的抑制作用(预占)制约。大型语言模型同样基于使用,其语言结构通过海量文本训练获得。本文测试了固有化与预占这两种对立统计力量是否同样在大模型中促进并限制语言生产力。实验表明,不同架构的大模型均能识别并复现生造词的构造性产出(固有化),在强制语境中实现非典型语义解读。然而,即使最大模型也无法将否定证据推广至新语言,统计上的预占并未使模型避免对语义合理但训练数据中从未出现的模式进行过度泛化。

原文摘要 · Abstract (English)

Usage-based theories of grammars posit that creative productivity of the structures of language is both bolstered and constrained by two distinct frequency signals: entrenchment, stemming from high frequency usage, and preemption, stemming from having never observed a particular linguistic structure in a context where one might expect that structure to appear. Large Language Models are also usage-based, in the sense that the structures of language are learned through exposure to vast amounts of text. Here, we test whether or not the opposing statistical forces of entrenchment and preemption also encourage and constrain linguistic productivity in LLMs. We demonstrate across model architectures that larger models recognize and can reproduce with nonce words constructional productivity (entrenchment) in cases of coercion, wherein the broader constructional context coerces an atypical interpretation of a lexical item. However, we also show that even the largest models do not extend negative evidence to novel language, and statistical preemption does not enable models to avoid overgeneralization of patterns that are semantically felicitous, but never observed in data.

语言模型生成能力认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。