arXiv:2508.19851cs.AI2025-08被引 4

用国际象棋评估大模型对世界状态的追踪能力,发现其长期推理存在缺陷。

Tracking World States with Language Models: State-Based Evaluation Using Chess

  • 以棋局状态为基准,通过合法走法分布评估模型理解力。
  • 实验显示模型在长序列中难以保持状态一致性,存在语义丢失。
  • 无需访问模型内部参数,适用于各类符号系统评测。

大型语言模型在结构化领域展现出涌现能力,暗示其可能隐式内化了高保真世界模型表征。尽管探测技术已在科学和游戏场景中揭示此类迹象,但这些方法依赖模型特定的内部激活,限制了可解释性和泛化性。本文提出一种基于国际象棋的、无需模型内部信息的通用状态评估框架,通过分析下游合法走法分布(状态可行动作)来估计预测与实际棋局状态之间的语义保真度。该方法比传统字符串匹配指标更贴合国际象棋的战略性与规则性,能有效捕捉状态追踪中的缺陷。实验表明,现有模型在长序列中难以维持一致的内部状态,暴露出结构化推理的局限。本框架为评估大模型的结构化推理能力提供了可靠工具,并可推广至多种符号环境。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit emergent capabilities in structured domains, suggesting they may implicitly internalize high-fidelity representations of world models. While probing techniques have shown promising signs of this in scientific and game-based settings, they rely on model-specific internal activations, which limit interpretability and generalizability. In this work, we propose a model-agnostic, state-based evaluation framework using chess as a benchmark to assess whether LLMs preserve the semantics of structured environments. Our method analyzes the downstream legal move distributions (state affordances) to estimate semantic fidelity between predicted and actual game states. This approach offers a more meaningful evaluation than conventional string-based metrics by aligning more closely with the strategic and rule-governed nature of chess. Experimental results demonstrate that our metrics capture deficiencies in state-tracking, highlighting limitations of LLMs in maintaining coherent internal models over long sequences. Our framework provides a robust tool for evaluating structured reasoning in LLMs without requiring internal model access, and generalizes to a wide class of symbolic environments.

大模型评估结构化推理国际象棋状态追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。