arXiv:2605.28003cs.CL2026-05被引 2

构建1.4万道前沿数学题数据集,验证大模型自主解题能力

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

论文配图:ResearchMath-14K: Scaling Research-Level Mathematics via Agents
图 1 · 摘自论文原文
  • 用多智能体流水线收集1.4万道研究级数学题
  • 新模型生成的参考文献量是旧版的5.6倍且虚假引用更多
  • 过滤后的错误解题路径仍可有效提升模型表现

数学前沿问题尚无已知解,但语言模型能否在无人干预下有效应对仍不明确。主要障碍在于缺乏大规模研究级数学数据集。为此,我们提出ResearchMath-14k,一个由多智能体管道从学术来源整理的14,056道问题集合,为目前最大规模的研究级数学问题数据集。我们还生成了22万条教师推理轨迹,发现存在反复出现的回避行为,如不尝试或虚构参考文献。在八款开源模型中,新一代模型每条轨迹生成的参考文献量是旧版的5.6倍,虚假参考文献达5.0倍。经智能体过滤后,将Qwen3模型从4B微调至30B参数,在基准上平均提升9.2分。这表明即使缺乏完整正确推理路径,经过筛选的开放问题求解尝试也能提供有效监督。ResearchMath-14k已公开供后续研究使用。

原文摘要 · Abstract (English)

The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully engage with such problems without human intervention. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of $14{,}056$ problems curated from academic sources via a multi-agent pipeline, making it the largest collection of research-level mathematical problems to date. We further generate ResearchMath-Reasoning, $220$K teacher trajectories from two open models, where we observe recurring avoidance behaviors such as non-attempts and fabricated references. Interestingly, across eight open-weight models, newer generations produce $5.6\times$ more references and $5.0\times$ more fake references per trace. After agentic filtering of ResearchMath-Reasoning, fine-tuning Qwen3 models from 4B to 30B parameters improves over base models by $9.2$ points on average. This shows that filtered open-problem attempts can provide useful supervision even without fully correct reasoning traces. We make ResearchMath-14k publicly available for future works on research-level mathematical reasoning.

数学推理多智能体大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。