arXiv:2607.06820cs.AI2026-07中稿 · ICML

让大模型用SageMath计算,解决数学难题效率提升超9.7个百分点。

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

论文配图:Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics
图 1 · 摘自论文原文
  • 大模型结合SageMath进行推理,获得可验证的数学计算反馈。
  • 平均性能提升9.7个百分点,最高达27.8个百分点,缩小了开源与闭源模型差距。
  • 适合数学研究者探索自动化猜想发现,尤其适合计算数学方向。

人工智能在数学领域的进展主要集中在自动形式化和定理证明,而计算机代数系统(CAS)在大模型智能体工作流中的作用仍被忽视。本文提出一种类ReAct的智能体架构,融合大模型推理与SageMath的可验证计算反馈,并引入Context7获取最新文档支持。我们在模拟计算数学研究循环的RealMath基准上评估该架构,针对前沿模型进行测试。同时,我们对RealMath基准提出改进:引入多步后处理流程与多阶段验证机制,显著提升问题集的质量与可靠性。实验表明,所有测试模型在接入SageMath后均取得显著性能提升,平均增益达+9.7个百分点,个体增益范围为1.5~27.8个百分点,有效缩小了开放权重与封闭模型间的差距。Qwen 3.7-Max受益最大,GPT-5.5在工具启用配置中达到最高求解率75.2%且消耗最少令牌。结果表明,融合CAS的大模型智能体是辅助数学家开展计算探索的有前景方向,推动自动化猜想发现进程。项目代码已公开。

原文摘要 · Abstract (English)

Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation. We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a computational-mathematics research loop. We also propose a refinement to the RealMath benchmark by introducing a multi-step post-processing procedure and a multi-stage validation pipeline, both of which improve the quality and reliability of the extracted problem set. Our experiments reveal substantial performance gains from SageMath access across all evaluated models on +9.7~pp on average, the gains range from 1.5~pp to 27.8~pp and narrow the gap between open-weight and closed models. Qwen~3.7-Max benefits from SageMath the most, while GPT-5.5 achieves the highest solve rate of $75.2\%$ and the lowest token usage among tool-enabled configurations. Our findings suggest that CAS-augmented agents represent a promising direction for assisting mathematicians in computational exploration, and we believe that this work is a step towards automated conjecture discovery. The project repository is available online.

数学AI大模型计算数学SageMath

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。