arXiv:2608.26950cs.AIcs.CL2026-08

为大模型数学智能设计过程级评估框架,揭示其真实推理能力

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

论文配图:From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities
图 1 · 摘自论文原文
  • 构建基于原子能力的结构化评估体系,追踪解题全过程
  • 相同终答案模型间存在显著过程能力差异,验证评估必要性
  • 适合研究大模型推理机制与下一代数学智能代理的开发者

大型语言模型正从端到端数学推理转向集成代理智能。然而,现有数学评测仅关注最终答案,缺乏对过程错误或逻辑严谨性的诊断价值,难以指导模型向稳健代理演进。为此,我们提出一种过程级评测框架,将问题求解中的代理行为与可复用的数学原子能力分类体系对齐。设计涵盖文本与多模态场景的规划、动作、反馈任务组合,并通过自动化流水线合成高质量解题轨迹,利用受控的LLM重写生成细粒度标注。实验表明,即使终答案准确率相近,不同模型在代理能力维度上仍呈现显著差异。这证明过程级评估对理解大模型真实潜力及推动下一代数学代理发展至关重要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation provides limited diagnostic value for identifying process-level failures or rigorous logic, failing to guide the transformation of LLMs into robust agents. To bridge this gap, we present a process-level benchmark designed to evaluate the inherent agentic mathematical reasoning abilities of LLMs. Our framework aligns problem-solving agentic behaviors with a structured taxonomy of reusable mathematical atomic capabilities. We design a comprehensive suite of planning, action, and feedback tasks across both textual and multimodal contexts, supported by an automated pipeline that synthesizes high-quality trajectories and produces fine-grained annotations via controlled LLM rewriting. Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles. This demonstrates that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.

大模型评估数学推理代理智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。