arXiv:2609.06100cs.LGcs.AI2026-09

用可验证证据优化语言模型决策,提升科学推理与工具使用能力。

VERPO: Verified Evidence Regularized Policy Optimization

论文配图:VERPO: Verified Evidence Regularized Policy Optimization
图 1 · 摘自论文原文
  • 将证据视为策略修正建议,分离无证据恢复与带符号修正。
  • 在多个任务上显著提升性能,最高达56.57%准确率。
  • 适合需要精准推理和可控生成的研究者与开发者。

可验证结果奖励指导语言模型后训练,但序列级优势无法识别应保留或修改的词元级决策。基于证据的教师通过回放带有特权反馈的采样轨迹提供更密集的监督。然而盲目模仿可能传递不影响任务成功的格式或推理风格变化。我们提出VERPO框架,将证据视为策略修正提案,同时保留结果目标。该方法分离无证据参考恢复与带符号的词元级证据修正。费舍尔证据对比沿估计的证据存在方向削弱修正。停止的词元级ZPD控制器根据局部奖励一致性与费舍尔移动成本调整接受度,而参考通道保持独立。在五个科学推理与工具使用任务中,各主干模型的最佳变体平均得分均超过最强基线:Qwen3-4B从0.6826升至0.6857,Qwen3-8B从0.6895升至0.7058,Llama-3.2-1B从0.4751升至0.5657。

原文摘要 · Abstract (English)

Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evidence-free reference restoration from signed token-level evidence corrections. Fisher Evidence Contrast attenuates corrections along an estimated evidence-presence direction. A stopped token-wise ZPD controller scales acceptance according to local reward alignment and Fisher movement cost, while the reference channel remains independent of acceptance. Across five scientific-reasoning and tool-use tasks, the best variant on each backbone exceeds the strongest compared baseline in average score. The averages rise from 0.6826 to 0.6857 on Qwen3-4B, from 0.6895 to 0.7058 on Qwen3-8B, and from 0.4751 to 0.5657 on Llama-3.2-1B.

强化学习推理优化语言模型证据引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。