用虚构语言测试大模型能否真正理解语法规则,发现其远不如人类。
The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang
- 设计虚构语言Camlang,通过语法规则和词典学习来检验推理能力。
- GPT-5在Camlang上准确率仅47%,远低于人类的87%。
- 揭示当前模型多依赖表面词汇匹配,缺乏系统性语法掌握。
大型语言模型在多项基准测试中表现优异,但其成功是源于真实推理还是模式匹配尚不明确。从认知科学视角出发,一个有效的检验方式是模型能否通过显式元语言推理掌握一种陌生语言——这是人类学习第二语言时可稳定实现的能力。为此,本文构建了虚构语言Camlang,其包含自然但未出现过的特征组合,由语法书与双语词典构成,模拟成人二语学习过程,可分离形态句法、词汇语义与句级推理错误。人类实验表明,参与者仅凭这些资源即可掌握并完成任务。我们基于CommonsenseQA构建了首个任务——Camlang-CSQA-v0,要求运用语法规则与词义映射作答。实验显示,GPT-5在英文上达98%精确率(EM),但在Camlang上仅为47%,显著低于人类的87%;其他先进推理模型表现更差。人工验证发现,多数模型成功源于浅层词汇对齐,而GPT-5虽有初步元语言意识,仍未实现系统性语法掌握。Camlang建立了一种具认知基础的评估范式,暴露了当前模型与人类元语言能力之间的根本差距。
原文摘要 · Abstract (English)
Large Language Models (LLMs) achieve gold-medal performance across many benchmarks, yet it remains unclear whether such success reflects genuine reasoning or pattern matching. From a cognitive science perspective, an informative test is whether models can master an unfamiliar language through explicit metalinguistic deductive learning, a paradigm where human learners can reliably internalise grammatical systems through metalinguistic reasoning. We address this question with Camlang, a novel constructed language that exhibits naturalistic yet unattested feature combinations. Camlang consists of two explicit resources, a grammar book and a bilingual dictionary, which mirror adult second-language learning via explicit grammar rules and lexical lookup, and enable us to disentangle errors in morpho-syntax, lexical semantics, and sentence-level reasoning. Human experiments show that these resources are sufficient for participants to acquire Camlang and successfully solve Camlang tasks. To operationalise evaluation, we adapt CommonsenseQA into Camlang, creating Camlang-CSQA-v0, the first task in a broader suite where solving questions requires applying grammar rules and lexical mappings. Experimental results show that GPT-5 achieves 98\% EM accuracy in English but only 47\% in Camlang, far below human performance at 87\%, while other state-of-the-art reasoning LLMs perform even worse. Human verification further reveals that most model successes stem from shallow lexical alignment while GPT-5 shows emerging metalinguistic awareness to a limited extent but not systematic grammatical mastery as humans. Camlang establishes a cognitively grounded evaluation paradigm that exposes fundamental gaps between current models and human metalinguistic competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。