RealMath用真实科研数学题评估大模型,发现其表现远超竞赛题
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
- 从科研论文和论坛直接构建真实数学任务集
- 多模型测试显示对研究级数学有意外强的处理能力
- 持续更新数据集防污染,适合数学研究辅助场景
现有大语言模型数学推理评测主要依赖竞赛题、形式证明或人工构造难题,无法反映真实科研中的数学实践。我们提出RealMath,一个直接来源于科研论文和数学论坛的新型基准,用于评估模型在真实数学任务上的表现。该方法解决了三大挑战:获取多样化的研究级内容、通过可验证命题实现可靠自动评估、设计可持续更新的数据集以降低污染风险。多模型实验结果表明,模型在研究级数学任务上的表现显著优于竞赛题场景,暗示当前模型已可作为数学工作者的有力助手,尽管在极难问题上仍存局限。RealMath代码与数据集已公开。
原文摘要 · Abstract (English)
Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions -- failing to capture the nature of mathematics encountered in actual research environments. We introduce RealMath, a novel benchmark derived directly from research papers and mathematical forums that assesses LLMs' abilities on authentic mathematical tasks. Our approach addresses three critical challenges: sourcing diverse research-level content, enabling reliable automated evaluation through verifiable statements, and designing a continually refreshable dataset to mitigate contamination risks. Experimental results across multiple LLMs reveal surprising capabilities in handling research mathematics compared to competition problems, suggesting current models may already serve as valuable assistants for working mathematicians despite limitations on highly challenging problems. The code and dataset for RealMath are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。