让模型自检更准,能发现两个错误答案的陷阱。
Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving
- 用强化学习训练双答案对比验证器,输出可信度和解释
- 在最佳N选一和自我反思策略中提升准确率与效率
- 能识别双错情况,适合需要高可靠性的推理任务
先进的测试时计算策略对扩展推理模型至关重要,但其效果受限于模型自身评估能力不足。本文提出一种基于强化学习(GRPO)的成对解释性验证器,可生成校准后的置信度分数及自然语言推理过程,用于评估生成解的可靠性。该验证器显著提升了最佳N选一、自我反思等测试时策略的准确性与效率。关键优势在于能有效识别复杂失败模式,例如当多个候选解均错误且表现一致时,仍能正确判断,而传统多数投票方法在此类场景下会失效。
原文摘要 · Abstract (English)
Advanced test-time computing strategies are essential for scaling reasoning models, but their effectiveness is capped by the models' poor self-evaluation. We propose a pairwise Explanatory Verifier, trained via reinforcement learning (GRPO), that produces calibrated confidence scores and associated natural language reasoning for generated solutions. Our verifier improves the accuracy and efficiency of test-time strategies like best-of-n and self-reflection. Crucially, it excels at identifying challenging failure modes, such as when both candidate solutions are identically incorrect, succeeding where standard methods like majority voting fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。