构建通用推理评估基准,检验模型在多领域错误检测能力。
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models

- 设计跨科学与逻辑九子领域的过程级评测基准
- 发现现有模型在数学外领域错误检测能力显著下降
- 揭示PRM不擅长知识类错误,LLM难检计算类错误
当前过程奖励模型(PRMs)在测试时扩展中展现出巨大潜力。由于大语言模型(LLMs)在处理广泛推理与决策任务时常产生错误的中间推理步骤,因此需要具备在真实场景中识别过程级错误的能力。然而,现有基准主要聚焦于数学推理,难以全面评估PRMs在多样化推理场景中的错误检测能力。为弥补这一差距,我们提出GR-Ben,一个专为评估PRMs在两大核心推理领域(科学与逻辑)及九个子领域表现而设计的过程级基准。我们在22种不同模型上进行大规模实验,涵盖PRMs与LLMs,得出两个关键发现:(1) 在数学推理以外的领域,现有PRMs和LLMs的错误检测能力明显较弱;(2) 总体而言,PRMs在识别知识类错误方面表现较差,而LLMs则在检测计算类错误方面能力不足。我们希望GR-Ben能推动面向通用领域的PRM研究,从而提升LLMs的推理能力。
原文摘要 · Abstract (English)
Currently, process reward models (PRMs) have exhibited remarkable potential for test-time scaling. Since large language models (LLMs) regularly generate flawed intermediate reasoning steps when tackling a broad spectrum of reasoning and decision-making tasks, PRMs are required to possess capabilities for detecting process-level errors in real-world scenarios. However, existing benchmarks primarily focus on mathematical reasoning, thereby failing to comprehensively evaluate the error detection ability of PRMs across diverse reasoning scenarios. To mitigate this gap, we introduce GR-Ben, a process-level benchmark specifically designed for assessing PRM's performance across two primary reasoning domains (science and logic) and nine subdomains. We conduct extensive experiments on a diverse set of 22 models, encompassing both PRMs and LLMs, and derive two key findings: (1) In domains beyond mathematical reasoning, the error-detection ability of existing PRMs and LLMs is found to be markedly weaker by comparison.(2) In general, PRMs are less adept at identifying knowledge-based errors, whereas LLMs exhibit poorer performance in detecting computational errors. We hope GR-Ben can foster future researches on PRMs for general domains, thereby enhancing the reasoning capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。