构建细粒度评估框架,检验推理模型对中间步骤错误的检测能力
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
- 设计6216个问题与83456个步骤标签,多维度评估推理过程
- 15个模型测试显示现有推理评分模型存在显著缺陷
- 适合研究复杂推理、模型评估与批判性思维的学者使用
过程级奖励模型(PRMs)在复杂推理与决策任务中至关重要,因每个中间步骤均影响整体推理质量。由于语言模型在推理过程中易出现各类隐性错误,当前的评估基准主要关注步骤正确性,缺乏对细微错误检测能力的系统性评测。为此,我们提出PRMBench,一个专为评估PRM细粒度错误识别能力而设计的过程级基准。该基准包含6,216个精心设计的问题和83,456个步骤级标签,从简洁性、合理性及敏感性等多维度进行评测。我们在15个模型上开展实验,涵盖开源与闭源大模型作为批评者模型,发现现有PRMs存在显著弱点。结果凸显过程级评估的挑战,并指明未来研究方向。我们希望PRMBench能成为推动PRM评估与发展的可靠基准。
原文摘要 · Abstract (English)
Process-level Reward Models (PRMs) are crucial for complex reasoning and decision-making tasks, where each intermediate step plays an important role in the reasoning process. Since language models are prone to various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios. However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically. To address this gap, we introduce PRMBench, a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs. PRMBench comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity. In our experiments on 15 models, spanning both open-source PRMs and closed-source large language models prompted as critic models, we uncover significant weaknesses in current PRMs. These findings underscore the challenges inherent in process-level evaluation and highlight key directions for future research. We hope PRMBench can be a robust bench for advancing research on PRM evaluation and development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。