研究大模型验证为何有效,发现跨模型验证效果更好。
When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers
- 跨模型家族验证比同族验证更有效
- 推理后训练提升跨模型改进能力
- 数学逻辑题最受益于验证机制
大语言模型既可作为问题求解器,也可充当解题验证者,从求解器生成的多个候选答案中筛选高质量结果。本文系统研究了在何种条件下验证能带来收益。通过在37个覆盖逻辑推理、结构谜题、符号计算、数学、常识、事实回忆和领域知识的9个基准上评估不同模型家族、规模及基础与推理后训练版本,发现:1)跨模型家族验证优于自验证或同族验证,且验证收益随求解器与验证器相似度增加而降低;2)推理后训练削弱自提升能力,但增强跨家族改进效果;3)数学与逻辑类任务更易通过验证获益。
原文摘要 · Abstract (English)
Large language models (LLMs) can act as both problem solvers and solution verifiers, where the latter select high-quality answers from a pool of solver-generated candidates. This raises the question of under what conditions verification pays off in solver-verifier systems. Prior work has conducted only limited studies of the factors influencing verification performance, focusing primarily on self-verification and examining neither the relationship between solver and verifier model families nor the effects of reasoning post-training. To rectify this, we present a systematic study across 37 models spanning multiple families, sizes, and base vs. post-trained variants, evaluated on 9 benchmarks covering logical reasoning, structured puzzles, symbolic computation, mathematics, commonsense, factual recall, and domain knowledge. In order to support our analysis, we introduce and empirically validate verifier gain, a metric that predicts the performance improvements from test-time verifier-based rejection sampling. Our experiments find that 1) verification across model families is more effective than either self-verification or verification within the same family, and more generally that the benefits of verification decrease as the solver and verifier become more similar, 2) reasoning post-training weakens self-improvement abilities but strengthens cross-family improvement, and 3) some tasks are inherently more amenable to improvement through verification, particularly mathematical and logical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。