测试大模型能否学习不可能语言,挑战语言习得的哲学基础
Large Language Models and Impossible Language Acquisition: "False Promise" or an Overturn of our Current Perspective towards AI
- 用语法反转和词数否定构造不可能语言,对比模型学习表现
- GPT-2小模型在自然语言上损失更低、收敛更快,反向句型损失高2.25倍
- 结果暗示大模型依赖统计模式而非人类式因果自纠,适合认知科学与AI哲学研究
在乔姆斯基极具争议的批评《CHATGPT的虚假承诺》中,大型语言模型(LLMs)被视为仅能预测模式,缺乏人类语言习得中的内在因果与自我修正机制,因此无法区分不可能语言。这构成了对人工智能理论基础的根本性挑战,整合了当前大模型方法论的核心问题,并体现了典型的先验理性主义视角。本文从语言学与心理学现有文献出发,结合一项实验,探究大模型在学习可能与不可能语言方面的能力。我们通过变换英语构造了一系列语法上不可能的语言,包括句子整体反转以及基于词数奇偶性的否定。分别在GPT-2 small模型与长短期记忆(LSTM)模型上进行了两轮受控实验。单次训练轨迹的描述性分析显示,相较于不可能语言条件,GPT-2 small模型在自然语言上的最终损失更低、收敛更快、困惑度更低;其中反转条件差异最大,损失比高达2.25倍。而LSTM模型在各条件下差异极小。由于实验为单次运行(每条件n=1),仅报告描述性比较,不进行正式统计推断。基于理论分析与描述性实证结果,本文提出在乔姆斯基理论框架内对大模型的新理解,并倡导从其“理性主义浪漫主义”范式转向功能主义与经验主义的研究范式。
原文摘要 · Abstract (English)
In Chomsky's provocative critique "The False Promise of CHATGPT," Large Language Models (LLMs) are characterized as mere pattern predictors that do not acquire languages via intrinsic causal and self-correction structures like humans, therefore are not able to distinguish impossible languages. It stands as a representative in a fundamental challenge to the intellectual foundations of AI, for it integrally synthesizes major issues in methodologies within LLMs and possesses an iconic a priori rationalist perspective. We examine this famous critique from both the perspective in pre-existing literature of linguistics and psychology as well as a research based on an experiment inquiring into the capacity of learning both possible and impossible languages among LLMs. We constructed a set of syntactically impossible languages by applying certain transformations to English. These include reversing whole sentences, and adding negation based on word-count parity. Two rounds of controlled experiments were each conducted on GPT-2 small models and long short-term memory (LSTM) models. Descriptive analysis of single-run training trajectories shows that GPT-2 small models exhibit lower final loss, faster convergence, and lower perplexity on natural language compared to impossible language conditions, with the reversed condition showing the largest departure (loss ratios up to 2.25 * natural). LSTM models, by contrast, show minimal differences across conditions. Given the single-run nature of our experiments (n=1 per condition), we report descriptive comparisons and caution that formal statistical inference is precluded. Based on theoretical analysis and descriptive empirical findings, we propose a new vision within Chomsky's theory towards LLMs, and a shift of theoretical paradigm outside Chomsky, from his "rationalist-romantics" paradigm to functionalism and empiricism in LLMs research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。