arXiv:2502.19058cs.CL2025-02被引 7

提出数学数据错误检测与诊断新基准,提升合成数据可靠性。

MathDebugger: Detecting and Diagnosing Errors in Synthetic Mathematical Data

  • 构建类型感知的错误检测框架,区分四类数学题错因
  • 14个大模型表现仍不足,细粒度错误识别能力弱
  • 错误类型提示可显著提升纠错效果,适合数据质检场景

合成数学数据已成为提升大语言模型推理能力的关键资源,但生成题目与解答中的错误会严重削弱其价值。我们提出MathDebugger,一个类型感知的基准,用于评估模型在合成数学数据中检测与诊断错误的能力。该基准包含2000个正确题目、2000个涵盖四类错误的错误题目,以及2000个标注解答,其中610个解答存在三类错误。所有实例均经人工验证并标注正确性,错误样本进一步细分为具体错误类型。人类标注一致性高,各类别Fleiss kappa值介于0.69至0.91之间。我们评估了14个代表性大语言模型和3个过程奖励模型,结果显示即使强推理模型也远未饱和,尤其在细粒度错误识别上表现有限。此外发现解题与验证能力存在系统性差距:专精数学或长文本推理的模型,并不必然优于通用模型进行数据审计。最后证明,明确的错误类型信息能提供可操作的修正指导,显著提升纠错效果。MathDebugger为开发更可靠的数学数据合成与质量控制流程提供了实用基准。

原文摘要 · Abstract (English)

Synthetic mathematical data has become an important resource for scaling the reasoning capabilities of large language models, yet errors in generated questions and solutions can substantially undermine its value. We introduce MathDebugger, a type-aware benchmark for evaluating whether models can detect and diagnose errors in synthetic mathematical data. MathDebugger contains 2,000 correct questions, 2,000 erroneous questions balanced across four error types, and 2,000 annotated solutions, including 610 erroneous solutions spanning three error types. Each instance is manually verified and labeled for correctness, with erroneous instances further assigned a fine-grained error category. Human annotation achieves substantial to near-perfect agreement, with per-type Fleiss kappa ranging from 0.69 to 0.91. We evaluate 14 representative large language models and three process reward models. The results show that even strong reasoning models remain far from saturating MathDebugger, particularly when identifying fine-grained error types. We also uncover a consistent solving-verification gap: models specialized for mathematical reasoning or long-form reasoning do not necessarily outperform their general-purpose counterparts when auditing mathematical data. Finally, we show that explicit error-type information provides actionable guidance for correcting erroneous questions and solutions, yielding consistent and statistically significant improvements. MathDebugger provides a practical benchmark for developing more reliable mathematical data synthesis and quality-control pipelines.

数学推理数据质检错误诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。