评测大模型推理过程的纠错能力,发现现有模型在不同思考方式下表现不一。
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
- 构建六类系统性推理模式的评测集,覆盖分解、整合等关键思维路径
- 包含2995条含错误的推理链,揭示当前纠错模型普遍存在缺陷
- 适合研究推理验证、智能体评估与大模型可信性的研究人员使用
过程奖励模型(PRMs)在复杂推理与问题求解任务中至关重要,通过验证每个中间推理步骤的正确性来支持长周期决策的LLM智能体。现实场景中,大模型可能采用多种推理模式(如分解、转换)解决问题,但不同模式下易产生错误。因此,PRMs需具备在各类推理模式下识别错误的能力。然而,现有评测基准主要关注步骤正确性,缺乏对多种推理模式的系统性评估。为此,我们提出Socratic-PRMBench,一个涵盖六种推理模式(变换、分解、重组、演绎、验证、整合)的新基准,包含2995条含缺陷的推理路径。实验表明,现有PRMs及提示为批评者的大模型均存在显著不足,暴露出当前模型在多样化推理模式下的评估能力薄弱。该基准旨在成为系统评估PRMs在多元推理情境下的综合性测试平台,推动其未来发展。
原文摘要 · Abstract (English)
Process Reward Models (PRMs) are crucial in complex reasoning and problem-solving tasks (e.g., LLM agents with long-horizon decision-making) by verifying the correctness of each intermediate reasoning step. In real-world scenarios, LLMs may apply various reasoning patterns (e.g., decomposition) to solve a problem, potentially suffering from errors under various reasoning patterns. Therefore, PRMs are required to identify errors under various reasoning patterns during the reasoning process. However, existing benchmarks mainly focus on evaluating PRMs with stepwise correctness, ignoring a systematic evaluation of PRMs under various reasoning patterns. To mitigate this gap, we introduce Socratic-PRMBench, a new benchmark to evaluate PRMs systematically under six reasoning patterns, including Transformation, Decomposition, Regather, Deduction, Verification, and Integration. Socratic-PRMBench}comprises 2995 reasoning paths with flaws within the aforementioned six reasoning patterns. Through our experiments on both PRMs and LLMs prompted as critic models, we identify notable deficiencies in existing PRMs. These observations underscore the significant weakness of current PRMs in conducting evaluations on reasoning steps under various reasoning patterns. We hope Socratic-PRMBench can serve as a comprehensive testbed for systematic evaluation of PRMs under diverse reasoning patterns and pave the way for future development of PRMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。