arXiv:2503.14495cs.CLcs.AI2025-03EMNLP被引 9

通过迭代反思提升大模型数学推理错误识别准确率

Temporal Consistency for LLM Reasoning Process Error Identification

  • 利用多轮自我反思的时序一致性改进验证判断
  • 7B/8B模型经优化后超越70B/72B模型及GPT-4o
  • 适合关注模型轻量化与推理可信度的研究者

验证对有效数学推理至关重要。本文提出一种新的时序一致性方法,验证器通过迭代修正前一轮评估结果来优化判断。与单轮验证或多模型辩论不同,该方法借助一系列自我反思行为的时序一致性提升验证准确率。在多个数学过程错误识别基准(Mathcheck、ProcessBench 和 PRM800K)上的实证评估显示,该方法持续优于基线。应用于近期 DeepSeek R1 精炼模型时,表现显著:7B/8B 模型在 ProcessBench 上超越所有 70B/72B 模型及 GPT-4o;14B 精炼模型经优化后性能接近 Deepseek-R1。代码已开源。

原文摘要 · Abstract (English)

Verification is crucial for effective mathematical reasoning. We present a new temporal consistency method where verifiers iteratively refine their judgments based on the previous assessment. Unlike one-round verification or multi-model debate approaches, our method leverages consistency in a sequence of self-reflection actions to improve verification accuracy. Empirical evaluations across diverse mathematical process error identification benchmarks (Mathcheck, ProcessBench, and PRM800K) show consistent performance improvements over baseline methods. When applied to the recent DeepSeek R1 distilled models, our method demonstrates strong performance, enabling 7B/8B distilled models to outperform all 70B/72B models and GPT-4o on ProcessBench. Notably, the distilled 14B model with our method achieves performance comparable to Deepseek-R1. Our codes are available at https://github.com/jcguo123/Temporal-Consistency

推理验证时序一致性模型精炼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。