arXiv:2608.11941cs.AI2026-08被引 1

语言模型可自主解决数学猜想,成本仅几十美元。

OEIS Open: How many conjectures can language models turn into theorems?

论文配图:OEIS Open: How many conjectures can language models turn into theorems?
图 1 · 摘自论文原文
  • 构建开源基准OEIS Open,用492个未解猜想测试模型
  • 模型以50美元预算解决147个猜想,准确率达30%
  • 即使接入数万篇论文,性能仍未提升,适合研究智能数学

我们构建了基于OEIS的492个开放数学猜想的基准OEIS Open,这些猜想由Tsoukalas等人在Lean中形式化。此前仅通过定制代理尝试过,而我们的开源评估代码可让任意通用语言模型(LM)参与测试,并防范模型作弊。结果显示,配备最少工具的模型在每轮50美元预算下解决了147个猜想,得分30%。OEIS Open Lite是随机选取的100个猜想子集,用于低成本评估;在每轮200美元预算下,当前最佳模型在该子集上得分44%。将模型接入来自arXiv的476,000篇数学文献并未提升其在OEIS Open Lite上的表现,使用更复杂的代理循环也无帮助。这些猜想数学意义尚不确定,多数此前关注极少。尽管如此,结果表明语言模型可在低预算下自主解决前沿数学猜想。

原文摘要 · Abstract (English)

We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code runs any generic language model (LM) against them, and is secure against LM cheating attempts. We find that LMs equipped with a minimal set of tools resolve 147 of these conjectures with a budget of \$50 per attempt, scoring 30% on OEIS Open. OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation. When evaluated with a budget of \$200 per attempt, the best current LM scores 44% on OEIS Open Lite. Giving LMs access to the mathematics literature via 476,000 papers from arXiv did not increase performance on OEIS Open Lite, and nor did using more sophisticated agent loops. The conjectures covered in this work are of uncertain mathematical significance, and most have likely received little previous attention. Nevertheless, our results show that LMs can resolve open research conjectures autonomously and at modest cost.

数学推理语言模型自动证明开源基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。