arXiv:2601.16823cs.CLcs.AI2026-01被引 1

用国际象棋测试大模型,发现其依赖记忆而非真正推理。

Disentangling generalization and memorization in large language models using chess

  • 用象棋位置密度区分记忆与泛化能力,无需了解训练数据。
  • 无相关先验时模型性能退化至随机水平,新模型进步也放缓。
  • 推理增强有效但边际收益递减,提示当前模型泛化力不足。

大型语言模型(LLMs)表现出惊人能力,但尚不清楚这些能力是源于精巧的回忆还是真正的推理。本文引入国际象棋作为受控测试平台,旨在分离这两种能力。利用棋局结构和可扩展的引擎评估,构建了从可靠记忆解决的常见局面到完全新颖需泛化的极端局面的位置分类体系。关键在于,该方法无需了解模型训练数据即可实现区分。结合对GPT系列的纵向分析及对Claude Opus、Gemini等当代模型的严格评估,结果揭示显著梯度:随着相关先验密度降低,性能持续下降。尤其在先验极少的任务中,基础模型表现退化至随机走子基准。尽管新模型有所改进,但在先验稀疏任务上进展明显放缓。此外,虽然推理增强型推断能提升性能,但其每令牌的相对增益在缺乏相关先验时递减。这表明系统性泛化存在局限,提示仅靠规模无法实现鲁棒性能,亟需超越规模的新机制。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genuine reasoning ability. We introduce chess as a controlled testbed aimed at disentangling these faculties. Leveraging the game's structure and scalable engine evaluations, we construct a taxonomy of positions varying in density of relevant priors - ranging from common states solvable by memorization to completely novel ones requiring generalization. Crucially, our approach achieves this distinction without requiring explicit knowledge of the models' training data. Applying this taxonomy, we combine a longitudinal analysis of the GPT lineage with a rigorous evaluation of contemporary models, including Claude Opus and Gemini. Our analysis reveals a steep gradient: performance consistently degrades as the density of relevant priors decreases. Notably, for tasks with few relevant priors, base model performance regresses to the random-play baseline. While newer models improve, progress slows significantly for tasks with sparse priors. Furthermore, while reasoning-augmented inference improves performance, its relative marginal benefit per token decreases in the absence of relevant priors. These results suggest limitations in systematic generalization, highlighting the need for mechanisms beyond scale to achieve robust performance when deprived of relevant priors.

大模型泛化能力推理测试象棋实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。