arXiv:2605.00677cs.LG2026-05AAAI

用混淆数论游戏测试大模型的真推理能力,发现专用模型更抗干扰。

Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game

  • 通过重命名符号构建无语义线索的数学环境,测试模型推理能力。
  • 通用模型在混淆后推理速度下降,准确率大幅降低;专用推理模型保持稳定。
  • 揭示了模型是否具备真正形式化推理能力,适合评估AI数学发现潜力。

尽管大语言模型在形式数学基准(如MiniF2F)上取得显著成果,但其表现究竟是源于真正的逻辑推理,还是对预训练数据中语义模式的匹配仍不明确。本文提出‘架构推理’——即在陌生数学领域仅依赖局部公理和定义合成形式化证明的能力,作为未来自动定理发现AI的必要能力。为此,我们引入混淆数论游戏(Obfuscated Natural Number Game),通过重命名Lean 4中自然数游戏的标识符,创建了一个零知识、封闭的测试环境。评估结果显示,所有先进模型均存在普遍延迟增加现象;性能鲁棒性出现分化:通用模型(Claude-Sonnet-4.5、GPT-4o)准确率显著下降,而推理专用模型(DeepSeek-R1、GPT-5、DeepSeek-Prover-V2)在缺乏语义提示的情况下仍保持原有准确率。该结果为评估数学推理的真实能力提供了可量化的指标。

原文摘要 · Abstract (English)

While Large Language Models have achieved notable success on formal mathematics benchmarks such as MiniF2F, it remains unclear whether these results stem from genuine logical reasoning or semantic pattern matching against pre-training data. This paper identifies Architectural Reasoning: the ability to synthesize formal proofs using exclusively local axioms and definitions within an alien math domain, as the necessary ability for future automated theorem discovery AI. We use the Obfuscated Natural Number Game, a benchmark to evaluate Architectural Reasoning. By renaming identifiers in the Natural Number Game in Lean 4, we created a zero-knowledge, closed environment. We evaluate state-of-the-art models, finding a universal latency tax where obfuscation increases inference time. The results also reveal a divergence in robustness: while general models (Claude-Sonnet-4.5, GPT-4o) suffer performance degradation, reasoning models (DeepSeek-R1, GPT-5, DeepSeek-Prover-V2) maintain the same accuracy despite the absence of semantic cues. These findings provide a quantitative metric for assessing the true capacity for mathematical reasoning.

大模型推理形式证明数学能力评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。