让大模型学会自我验证,能显著提升推理能力与效率。
Learning to Self-Verify Makes Language Models Better Reasoners
- 通过多任务强化学习,同时训练生成与自检能力。
- 自检能力提升后,推理准确率接近标准训练水平。
- 适合需要高可靠性推理的场景,如数学和逻辑任务。
近期大型语言模型在复杂任务中生成有希望的推理路径方面表现出色。然而,尽管生成能力强,它们在自我验证答案方面仍表现薄弱,反映出生成与自检能力之间的持续不对称。本文深入研究了这一不对称性在整个训练过程中的演变,发现即使在同一任务上,提升生成能力也不会带来相应的自检能力提升。有趣的是,反过来,学习自检能有效改善生成性能,在达到与标准生成训练相当的准确率的同时,产生更高效、更有效的推理轨迹。基于此观察,我们进一步探索将自检融入生成训练的方法,提出一种多任务强化学习框架,将生成与自检作为两个独立但互补的目标进行优化。在多个基准和模型上的大量实验表明,相较于仅训练生成,该方法在生成和自检能力上均取得显著提升。
原文摘要 · Abstract (English)
Recent large language models (LLMs) achieve strong performance in generating promising reasoning paths for complex tasks. However, despite powerful generation ability, LLMs remain weak at verifying their own answers, revealing a persistent capability asymmetry between generation and self-verification. In this work, we conduct an in-depth investigation of this asymmetry throughout training evolution and show that, even on the same task, improving generation does not lead to corresponding improvements in self-verification. Interestingly, we find that the reverse direction of this asymmetry behaves differently: learning to self-verify can effectively improve generation performance, achieving accuracy comparable to standard generation training while yielding more efficient and effective reasoning traces. Building on this observation, we further explore integrating self-verification into generation training by formulating a multi-task reinforcement learning framework, where generation and self-verification are optimized as two independent but complementary objectives. Extensive experiments across benchmarks and models demonstrate performance gains over generation-only training in both generation and verification capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。