arXiv:2509.18383cs.AIcs.DM2025-09被引 9

测试大模型能否证明简单新猜想,发现其有推理能力但跨文献整合仍弱。

Gödel Test: Can Large Language Models Solve Easy Conjectures?

  • 设计戈德尔测试:让模型证明五个组合优化中的未解简单猜想。
  • 在三个简单问题上接近正确,一个问题甚至推翻原猜想并给出有效解。
  • 跨论文综合推理失败,说明模型仍缺乏深层知识整合能力。

前沿大模型在高中与本科数学竞赛中表现突出,但尚不清楚它们能否解决更高级数学领域中的新且简单的猜想。本文提出戈德尔测试:评估模型能否为尚未解决的简单猜想提供正确证明。我们研究了GPT-5在五个组合优化猜想上的表现,每个问题均来自一到两篇原始论文,我们隐藏了自身提出的猜想,并对模型推理过程进行详细分析。在三个较易问题上,GPT-5生成了近乎正确的解法;在第二个问题中,模型推导出一个不同近似保证,经验证后否定了原猜想,同时提供了有效解法。在第四个问题(需结合两篇论文结果)上失败。第五个问题更难,无已验证猜想,模型提出了与我们相同的算法,但在分析阶段失败,表明证明难度超出预期。尽管样本量小,结果表明模型在常规推理上已有进展,偶现原创性,但在跨文献整合方面仍存在明显局限。GPT-5或标志着迈向最终通过戈德尔测试的早期一步。

原文摘要 · Abstract (English)

Recent announcements from frontier AI model labs have highlighted strong results on high-school and undergraduate math competitions. Yet it remains unclear whether large language models can solve new, simple conjectures in more advanced areas of mathematics. We propose the Gödel Test: evaluating whether a model can produce correct proofs for very simple, previously unsolved conjectures. To this end, we study the performance of GPT-5 on five conjectures in combinatorial optimization. For each problem, we provided one or two source papers from which the conjecture arose, withheld our own conjecture, and then assessed the model's reasoning in detail. On the three easier problems, GPT-5 produced nearly correct solutions; for Problem 2 it even derived a different approximation guarantee that, upon checking, refuted our conjecture while providing a valid solution. The model failed on Problem 4, which required combining results from two papers. On Problem 5, a harder case without a validated conjecture, GPT-5 proposed the same algorithm we had in mind but failed in the analysis, suggesting the proof is more challenging than expected. Although our sample is small, the results point to meaningful progress on routine reasoning, occasional flashes of originality, and clear limitations when cross-paper synthesis is required. GPT-5 may represent an early step toward frontier models eventually passing the Gödel Test.

大模型数学证明推理能力戈德尔测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。