arXiv:2605.09012cs.AI2026-05

评测大模型在数学论文中精准检索证明所需工具的能力

Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

论文配图:Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics
图 1 · 摘自论文原文
  • 构建从部分证明中检索数学工具的基准,支持精确匹配与上下文验证
  • 当前最佳模型仅7.0%正确识别适用定理,显示其难以判断工具是否可用
  • 适合研究数学推理、文献检索或AI辅助证明的学者使用

大型语言模型在封闭世界数学推理中表现日益出色,但科研辅助还需基于文献的工具检索能力。当证明进入非平凡步骤时,理想助手应判断所需工具(如引理)是否存在,定位合适文献,并验证其假设与当前证明上下文一致。为此,我们提出Re$^2$Math,一个面向部分数学证明的工具接地检索基准。每个实例基于主定理证明中的候选引用构建,包含层级化上下文和可选泄漏控制提示。任务设计为来源接地但引用无偏,只要满足证明需求的定理即可接受。评估采用发布冻结的检索结果,确保可复现性;基准本身支持自动持续扩展。在当前测试集上,最优固定评判工具准确率仅为7.0%,尽管源定位率较高,表明现有系统常能检索到有效命题,却无法确认其适用于具体证明步骤。通过解耦引用召回、源定位与证明间隙充分性,Re$^2$Math将文献驱动的数学工具使用转化为可控诊断任务。

原文摘要 · Abstract (English)

Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the literature. When a proof reaches a non-trivial step, a useful assistant should determine whether the needed tool (e.g., a lemma) already exists, identify a suitable scholarly source, and verify that its assumptions align with the current proof context. To rigorously evaluate such capabilities, we introduce Re$^2$Math, a benchmark for tool-grounded retrieval from partial mathematical proofs. Each instance is built from a candidate instrumental citation in the proof of a main theorem, with hierarchical context and an optional leakage-controlled anchor hint. We also make the task source-grounded yet citation-agnostic in that any admissible theorem sufficient for the proof transition is accepted. Evaluation uses a release-frozen retrieval artifact, ensuring reproducibility, while the benchmark itself supports automatic, continual expansion with newly constructed instances. On the current benchmark test set, the best fixed-judge ToolAcc reaches 7.0%, despite substantially higher rates of source grounding, indicating that current systems often retrieve valid statements but fail to establish their applicability to the local proof step. By decoupling citation recall, grounding, and proof-gap sufficiency, Re$^2$Math transforms literature-grounded mathematical tool use into a controlled diagnostic task.

数学推理文献检索基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。