arXiv:2503.10573cs.LG2025-03被引 7

对比8个大模型的数学推理能力,发现深求模型表现突出。

Evaluating Mathematical Reasoning Across Large Language Models: A Fine-Grained Approach

  • 用三个独立数据集系统评估8个主流大模型的数学推理能力。
  • 深求R1在多数领域表现优异,正式逻辑任务准确率最高。
  • 小模型蒸馏后性能显著下降,Gemini 2.0 Flash响应最快速。

随着人工智能快速发展,大语言模型在医疗、工程、科学、教育及数学推理等领域影响深远。其中,数学推理因需多步逻辑与抽象泛化而尤为挑战。尽管已有研究探讨大模型在推理任务中的表现,但覆盖模型家族广泛性与深度的综合评估仍有限。本研究系统评估了包括两个最新深求模型在内的八款领先大模型,采用三个独立基准数据集进行测试。结果表明:(1) 深求R1在多数领域表现媲美o1,且在MMLU Formal Logic基准上准确率最高;(2) 蒸馏版本如DeepSeek-1.5B出现显著性能下降;(3) Gemini 2.0 Flash响应延迟最低。此外,我们分析了架构设计、训练范式与优化策略对推理性能的影响。这些发现为当前大模型在数学领域的能力与局限提供了新见解,并为未来更符合严谨推理需求的模型开发提供指导。

原文摘要 · Abstract (English)

With the rapid advancement of Artificial Intelligence (AI), Large Language Models (LLMs) have significantly impacted a wide array of domains, including healthcare, engineering, science, education, and mathematical reasoning. Among these, mathematical reasoning remains a particularly challenging capability, often requiring multi-step logic and abstract generalization. While prior work has explored LLM performance on reasoning tasks, comprehensive evaluations that span both depth and breadth across model families remain limited. In this study, we present a systematic evaluation of mathematical reasoning abilities across eight leading LLMs, including two recent DeepSeek models, using three independent benchmark datasets. Our analyses reveal several key findings: (1) DeepSeek-R1 performs competitively with o1 across most domains and achieves the highest accuracy on the MMLU Formal Logic benchmark; (2) distilled variants, such as DeepSeek-1.5B, exhibit substantial performance degradation; and (3) Gemini 2.0 Flash achieves the lowest response latency. Beyond quantitative metrics, we explore how architectural choices, training paradigms, and optimization strategies contribute to variation in reasoning performance. These findings provide new insights into the capabilities and limitations of current LLMs in mathematical domains, and offer guidance for the development of future models better aligned with rigorous reasoning demands.

数学推理大模型评测DeepSeek性能分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。