大模型在数学推理上仍难理解现实场景中的问题表述。
From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics
- 构建新基准测试,将抽象数学题转化为真实情境下的问题
- 开放模型在情境中准确率下降34点,大模型也仅提升有限
- 正确理解问题比解题更重要,需结合训练与场景数据
大型语言模型在多项数学基准测试中已接近专家水平,但其在实际应用中的可靠性仍未提升。本文聚焦于上下文数学推理,即从描述性场景中提取数学核心。我们提出了ContextMATH基准,将AIME和MATH-500题目重构为两种情境:场景定位(SG),将抽象问题嵌入现实叙事而不增加推理复杂度;复杂度扩展(CS),将明确条件转化为子问题以模拟实际约束的呈现方式。评估61个专有与开源模型发现:平均而言,开源模型在SG和CS上分别下降13和34分,专有模型下降13和20分。错误分析显示,主要失误源于问题表述错误,且随原题难度上升而加剧。正确表述成为成功前提,且其有效性随模型规模提升,表明大模型在理解与推理上均有进步。然而,表述与推理仍是相互补充的双重瓶颈。最后,使用情景数据微调可提升性能,仅训练表述无效,但性能差距仍显著,凸显上下文数学推理是当前大模型的核心挑战。
原文摘要 · Abstract (English)
Large language models now solve many benchmark math problems at near-expert levels, yet this progress has not fully translated into reliable performance in real-world applications. We study this gap through contextual mathematical reasoning, where the mathematical core must be formulated from descriptive scenarios. We introduce ContextMATH, a benchmark that repurposes AIME and MATH-500 problems into two contextual settings: Scenario Grounding (SG), which embeds abstract problems into realistic narratives without increasing reasoning complexity, and Complexity Scaling (CS), which transforms explicit conditions into sub-problems to capture how constraints often appear in practice. Evaluating 61 proprietary and open-source models, we observe sharp drops: on average, open-source models decline by 13 and 34 points on SG and CS, while proprietary models drop by 13 and 20. Error analysis shows that errors are dominated by incorrect problem formulation, with formulation accuracy declining as original problem difficulty increases. Correct formulation emerges as a prerequisite for success, and its sufficiency improves with model scale, indicating that larger models advance in both understanding and reasoning. Nevertheless, formulation and reasoning remain two complementary bottlenecks that limit contextual mathematical problem solving. Finally, we find that fine-tuning with scenario data improves performance, whereas formulation-only training is ineffective. However, performance gaps are only partially alleviated, highlighting contextual mathematical reasoning as a central unsolved challenge for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。