LLM数学推理常出错,新评估方法能发现逻辑漏洞。
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
- 提出MAPLE评分法,综合判断推理错误、冗余和有效性
- 发现仅看答案正确率会忽略中间逻辑缺陷
- 适合研究大模型数学能力或评测系统设计者
大型语言模型(LLMs)在自然语言任务中展现出巨大潜力,但在数学推理方面仍面临显著挑战,尤其在执行精确的多步逻辑时。然而,当前评估框架仅基于最终答案的准确率来评判模型表现,这仅反映结果而非推理过程。本研究通过引入一种新型评估框架,揭示了此类评价方式的局限性。我们提出一种名为MAPLE的评分指标,通过整合错误率、冗余度和有效性,全面量化推理过程中的不一致程度。该方法能更真实地反映模型在复杂数学问题上的思维轨迹,揭示传统评估无法捕捉的深层缺陷。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。