从人类认知角度评估大模型的数学能力,发现现有方法高估了30%-40%。
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
- 按人类解题过程分三阶段设计9个细粒度维度评估
- 7个主流大模型真实数学能力被高估30%-40%
- 适合研究模型推理缺陷与提升方向的学者
尽管大语言模型在解决复杂数学任务上展现潜力,但现有评估仅依赖整体答案准确率,无法真实反映其能力。本文提出CogMath,从人类认知视角全面评估大模型数学能力。受心理学理论启发,将人类解题过程形式化为三个阶段:问题理解、求解过程和解答总结,并在此基础上考察数值计算、知识运用及反事实推理等视角,构建共9个细粒度评估维度。每个维度采用“提问-判断-参考”多智能体系统生成评估问题,只有在全部9个维度均表现优异的模型才被认为真正掌握该问题。在三个基准上的应用表明,7个主流大模型的数学能力被高估了30%-40%。同时定位其在各阶段与维度中的强弱项,为提升模型推理能力提供深入洞见。
原文摘要 · Abstract (English)
Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy, which are insufficient for assessing their authentic capabilities. In this paper, we propose \textbf{CogMath}, which comprehensively assesses LLMs' mathematical abilities through the lens of human cognition. Specifically, inspired by psychological theories, CogMath formalizes human reasoning process into 3 stages: \emph{problem comprehension}, \emph{problem solving}, and \emph{solution summarization}. Within these stages, we investigate perspectives such as numerical calculation, knowledge, and counterfactuals, and design a total of 9 fine-grained evaluation dimensions. In each dimension, we develop an ``\emph{Inquiry}-\emph{Judge}-\emph{Reference}'' multi-agent system to generate inquiries that assess LLMs' mastery from this dimension. An LLM is considered to truly master a problem only when excelling in all inquiries from the 9 dimensions. By applying CogMath on three benchmarks, we reveal that the mathematical capabilities of 7 mainstream LLMs are overestimated by 30\%-40\%. Moreover, we locate their strengths and weaknesses across specific stages/dimensions, offering in-depth insights to further enhance their reasoning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。