arXiv:2510.12829cs.CLcs.AI2025-10

用大模型协作证明数学题,五道奥数题全解,六十六个猜想解决近三分之一。

Mathematics with large language models as provers and verifiers

  • 多实例大模型协同推理,分角色扮演证明者与验证者。
  • 成功解决2025年国际数学奥赛五道题,破解66个数论猜想约三分之一。
  • 结果经Lean形式化验证与人工核对,防止幻觉误导。

2024至2025年,关于大语言模型在定理证明方面的能力讨论陆续出现令人瞩目的成果,涵盖难题(如国际数学奥林匹克竞赛题目)及为测试人工智能而设计的猜想。本文报告了利用GPT-5模型的不同证明者与验证者实例协作,通过特定协议实现的一次定理证明突破。为确保证明无幻觉,最终结果由Lean证明助手进行形式化验证,且前提与结论的逻辑一致性由人工确认。该方法虽不完整也不精确,但仍成功解决了2025年国际数学奥赛中的五道题目,并破解了科恩(Cohen, Journal of Integer Sequences, 2025)提出的六十六个数论猜想中的约三分之一。

原文摘要 · Abstract (English)

During 2024 and 2025 the discussion about the theorem-proving capabilities of large language models started reporting interesting success stories, mostly to do with difficult exercises (such as problems from the International Mathematical Olympiad), but also with conjectures [Feldman & Karbasi, arXiv:2509.18383v1] formulated for the purpose of verifying whether the artificial intelligence could prove it. In this paper we report a theorem proving feat achieved by ChatGPT by using a protocol involving different prover and verifier instances of the gpt-5 model working collaboratively. To make sure that the produced proofs do not suffer from hallucinations, the final proof is formally verified by the lean proof assistant, and the conformance of premises and conclusion of the lean code is verified by a human. Our methodology is by no means complete or exact. It was nonetheless able to solve five out of six 2025 IMO problems, and close about a third of the sixty-six number theory conjectures in [Cohen, Journal of Integer Sequences, 2025].

数学证明大模型形式验证奥赛题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。