评测大模型在数学创作题上的真实创新能力,发现现有模型表现有限。
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
- 构建新基准DeepMath-Creative,涵盖代数、几何等领域的创造性数学题
- 顶尖模型O3 Mini在宽松评分下仅70%正确率,复杂题几乎无解
- 揭示当前模型依赖记忆重组,缺乏真正数学创造能力,适合研究者参考
为提升大语言模型(LLM)的数学能力,DeepMath团队启动开源计划,致力于开发开放数学LLM并系统评估其数学创造力。尽管近期数学LLM多聚焦于推理能力,但其创造性能力仍受忽视,评估数据集稀缺。为此,本文提出数学创造力评估标准,引入全新的高质量基准DeepMath-Creative,涵盖代数、几何、分析等领域的构造性问题。我们对主流LLMs进行系统评估。结果表明,即使采用宽松评分标准(仅关注核心解题要素,忽略微小逻辑漏洞、不完整证明或冗余表达),最佳模型O3 Mini在基础本科级构造题上也仅达70%准确率,复杂问题表现急剧下降,无法提出有效策略应对开放问题。这表明当前模型在熟悉问题上表现出的解题能力可能源于记忆模式的重组,而非真正的创造性洞察或新颖整合。
原文摘要 · Abstract (English)
To advance the mathematical proficiency of large language models (LLMs), the DeepMath team has launched an open-source initiative aimed at developing an open mathematical LLM and systematically evaluating its mathematical creativity. This paper represents the initial contribution of this initiative. While recent developments in mathematical LLMs have predominantly emphasized reasoning skills, as evidenced by benchmarks on elementary to undergraduate-level mathematical tasks, the creative capabilities of these models have received comparatively little attention, and evaluation datasets remain scarce. To address this gap, we propose an evaluation criteria for mathematical creativity and introduce DeepMath-Creative, a novel, high-quality benchmark comprising constructive problems across algebra, geometry, analysis, and other domains. We conduct a systematic evaluation of mainstream LLMs' creative problem-solving abilities using this dataset. Experimental results show that even under lenient scoring criteria -- emphasizing core solution components and disregarding minor inaccuracies, such as small logical gaps, incomplete justifications, or redundant explanations -- the best-performing model, O3 Mini, achieves merely 70% accuracy, primarily on basic undergraduate-level constructive tasks. Performance declines sharply on more complex problems, with models failing to provide substantive strategies for open problems. These findings suggest that, although current LLMs display a degree of constructive proficiency on familiar and lower-difficulty problems, such performance is likely attributable to the recombination of memorized patterns rather than authentic creative insight or novel synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。