arXiv:2604.18584cs.AIcs.DL2026-04被引 6

构建首个覆盖47国17语言的数学奥赛基准,评估模型解题与检索能力

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

论文配图:MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
图 1 · 摘自论文原文
  • 整合3万+道奥赛题,涵盖20年赛事与多语言多领域
  • 顶尖模型解题准确率仅78.4%,检索模型难识别等价题目
  • 首次引入检索增强解题测试,优质检索可提升性能12%

数学问题求解仍是大模型推理能力的重要挑战,但现有基准在规模、语言覆盖和任务多样性上受限。我们提出MathNet,一个高质量、大规模、多模态、多语言的奥林匹克级数学问题数据集,以及用于评估生成模型数学推理能力和基于嵌入系统数学检索性能的基准。MathNet覆盖47个国家、17种语言,跨越二十年竞赛,包含30,676道专家编写的问题及解答,涵盖多样化领域。除核心数据集外,我们构建了由人工专家精心筛选的数学等价与结构相似问题对组成的检索基准。MathNet支持三项任务:(i) 问题求解,(ii) 数学感知检索,(iii) 检索增强型问题求解。实验表明,即使最先进的推理模型(Gemini-3.1-Pro达78.4%,GPT-5达69.3%)仍面临挑战,而嵌入模型在检索等价问题时表现不佳。我们进一步发现,检索增强生成性能高度依赖检索质量;例如,DeepSeek-V3.2-Speciale在检索优化后性能提升高达12%,取得基准最高分。MathNet是目前最大且最高质量的奥赛数据集,并提供首个数学问题检索评估基准,数据集与基准已公开发布于https://mathnet.mit.edu。

原文摘要 · Abstract (English)

Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a high-quality, large-scale, multimodal, and multilingual dataset of Olympiad-level math problems together with a benchmark for evaluating mathematical reasoning in generative models and mathematical retrieval in embedding-based systems. MathNet spans 47 countries, 17 languages, and two decades of competitions, comprising 30,676 expert-authored problems with solutions across diverse domains. In addition to the core dataset, we construct a retrieval benchmark consisting of mathematically equivalent and structurally similar problem pairs curated by human experts. MathNet supports three tasks: (i) Problem Solving, (ii) Math-Aware Retrieval, and (iii) Retrieval-Augmented Problem Solving. Experimental results show that even state-of-the-art reasoning models (78.4% for Gemini-3.1-Pro and 69.3% for GPT-5) remain challenged, while embedding models struggle to retrieve equivalent problems. We further show that retrieval-augmented generation performance is highly sensitive to retrieval quality; for example, DeepSeek-V3.2-Speciale achieves gains of up to 12%, obtaining the highest scores on the benchmark. MathNet provides the largest high-quality Olympiad dataset together with the first benchmark for evaluating mathematical problem retrieval, and we publicly release both the dataset and benchmark at https://mathnet.mit.edu.

数学推理多语言检索增强奥赛数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。