arXiv:2604.04791cs.CL2026-04

对比大模型与专家在数学建模中的表现,发现大模型执行环节存在明显短板。

How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling

  • 分阶段评估大模型在建模全流程的表现,用专家标准验证可靠性。
  • 大模型在问题识别上表现良好,但在求解、编程和分析阶段错误频发。
  • 缺陷源于需求不明确、缺少验证,且错误会跨阶段传播。

大语言模型在推理基准测试中表现优异,但其解决需要端到端流程的真实问题能力尚不明确。数学建模竞赛为评估此类端到端求解能力提供了严苛的测试环境。本文提出一种面向问题、分阶段的评估框架,基于专家验证标准评估大模型在建模各阶段的表现。通过对比中国研究生数学建模竞赛题目的自动评分与独立专家判断,验证了该框架的可靠性,其一致性显著优于现有评估方法。利用该框架,我们发现当前顶尖大模型存在理解-执行差距:在问题识别与建模等早期阶段表现良好,但在模型求解、代码实现和结果分析等执行阶段持续存在缺陷。这一差距即使在模型规模扩大后依然存在。进一步分析表明,失败源于需求描述不充分、缺乏验证和校验机制,错误在各阶段间传播且无法纠正。研究提示,突破该瓶颈需超越单纯模型扩容,为大模型应用于复杂现实问题提供重要参考。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved strong performance on reasoning benchmarks, yet their ability to solve real-world problems requiring end-to-end workflows remains unclear. Mathematical modeling competitions provide a stringent testbed for evaluating such end-to-end problem-solving capability. We propose a problem-oriented, stage-wise evaluation framework that assesses LLM performance across modeling stages using expert-verified criteria. We validate the framework's reliability by comparing automatic scores with independent human expert judgments on problems from the China Postgraduate Mathematical Contest in Modeling, demonstrating substantially stronger alignment than existing evaluation schemes. Using this framework, we reveal a comprehension-execution gap in state-of-the-art LLMs: while they perform well in early stages such as problem identification and formulation, they exhibit persistent deficiencies in execution-oriented stages including model solving, code implementation, and result analysis. These gaps persist even with increased model scale. We further trace these failures to insufficient specification, missing verification, and lack of validation, with errors propagating across stages without correction. Our findings suggest that bridging this gap requires approaches beyond model scaling, offering insights for applying LLMs to complex real-world problem solving.

大模型评估数学建模执行缺陷端到端推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。