arXiv:2606.29894cs.IRcs.AI2026-06中稿 · the 3rd AI for Mat…

首个无需人工标注的数学信息检索评估基准,专为测试AI解题时的查资料能力而设计。

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

  • 用大模型自动生成题目摘要和主题,构建自动化的数学检索任务
  • 在代数与微积分等符号密集领域,最强模型仍表现不佳
  • 适合研究数学AI、检索系统或评测工具的开发者使用

随着智能体系统处理更复杂的数学任务,它们越来越依赖信息检索(IR)来查询问题数据库、定理库和教育资源。然而,由于难以直接分离检索器对下游性能的影响,选择合适的检索器仍具挑战性。现有检索专用基准常无法捕捉细粒度的数学相关性,误罚相关文档。为此,我们提出SABER-Math,首个完全自动化、无需专家标注的数学信息检索评估基准。基于28.3万道高中数学题及其解答,该基准通过三步构建挑战性重排序任务:(i) 使用大模型提取每道题的简洁解题摘要和数学主题;(ii) 基于本体主题与解题摘要的相似性,发现每查询的相关文档;(iii) 采用瑞士式大模型偏好锦标赛生成细粒度相关性评分。我们评估了词法检索器、专用数学检索系统及最新嵌入模型。结果显示,尽管现代嵌入模型显著优于传统与数学专用基线,但在代数与微积分等符号密集领域仍表现受限。更重要的是,通用IR基准如MTEB无法可靠预测数学性能,尤其对最新嵌入模型,凸显了构建数学专用检索基准的必要性。

原文摘要 · Abstract (English)

As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources. However, choosing the right retriever remains difficult, as it is infeasible to directly isolate its effect on downstream performance. On the other hand, existing retrieval-specific benchmarks often fail to capture fine-grained mathematical relevance, penalizing relevant documents. We address this gap by introducing SABER-Math, the first fully automated benchmark for evaluating mathematical IR without expert annotation. Starting from 283K high-school-level math problems with solutions, SABER-Math builds challenging reranking tasks in three steps: (i) first, LLMs extract concise solution summaries and mathematical topics for each problem; (ii) then, per-query relevant documents are discovered using ontology topic-based and lexical solutions-summary-based similarities, and (iii) finally, a Swiss-style LLM preference tournament produces fine-grained relevance ratings for the documents. We evaluate lexical retrievers, specialized mathematical retrieval systems, and recent embedding models. We find that while modern embedding models substantially outperform classical and math-specific baselines, even the strongest systems struggle in symbol-heavy domains like Algebra and Calculus. Importantly, we show that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.

信息检索数学AI大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。