arXiv:2604.22597cs.AI2026-04被引 1

用大模型替代符号比对,让数学推理评估更可靠

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity

论文配图:Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity
图 1 · 摘自论文原文
  • 用大模型作裁判,自动判断答案对错
  • 在多个基准上超越传统符号比对方法
  • 适合研究数学推理与智能系统评估的人

大语言模型在数学推理任务上取得显著进展,常通过对比生成答案与标准答案来评估模型能力。现有方法依赖符号化数学比对,难以泛化到不同表达形式和解题格式。本文提出一种基于大模型的评估框架,可准确识别多样化的数学表达与答案形式。我们在Lighteval和SimpleRL两个主流框架中揭示了符号比对的失败案例,并证明本方法在准确性上明显优于常用方法。该框架提升了评估可靠性,有助于更精准地监控模型性能,推动数学求解与智能系统的发展。

原文摘要 · Abstract (English)

Recent advancements in large language models have led to significant improvements across various tasks, including mathematical reasoning, which is used to assess models' intelligence in logical reasoning and problem-solving. Models are evaluated on mathematical reasoning benchmarks by verifying the correctness of the final answer against a ground truth answer. A common approach for this verification is based on symbolic mathematics comparison, which fails to generalize across diverse mathematical representations and solution formats. In this work, we offer a robust and flexible alternative to rule-based symbolic mathematics comparison. We propose an LLM-based evaluation framework for evaluating model-generated answers, enabling accurate evaluation across diverse mathematical representations and answer formats. We present failure cases of symbolic evaluation in two popular frameworks, Lighteval and SimpleRL, and compare them to our approach, demonstrating clear improvements over commonly used methods. Our framework enables more reliable evaluation and benchmarking, leading to more accurate performance monitoring, which is important for advancing mathematical problem-solving and intelligent systems.

数学推理评估框架大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。