用强化学习让大模型自主查证信息,准确率提升30%。
Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning
- 通过在线强化学习训练大模型自主检索与推理
- 联合准确率最高提升30%,证据得分翻倍
- 适合需要高可信度信息验证的研究与应用
利用大语言模型(LLMs)进行主张验证近年来受到广泛关注,因其强大的推理能力与透明的验证过程优于传统仅回答式判断。然而,现有在线主张验证方法仍主要依赖提示工程或预设推理流程,缺乏统一训练以提升必要技能。为此,我们提出Veri-R1,一个在线强化学习(RL)框架,使大模型能够与搜索引擎交互,并接收显式奖励信号以塑造其规划、检索和推理行为。这种动态交互更贴近真实验证场景,培养了全面的验证能力。实证结果表明,Veri-R1在联合准确率上最高提升30%,证据得分翻倍,常优于更大规模模型。消融实验揭示了奖励组件的影响及输出逻辑值与标签准确性的关联。结果表明,在线强化学习对精确且忠实的主张验证具有显著有效性,为未来研究奠定基础。代码已开源,支持社区发展。
原文摘要 · Abstract (English)
Claim verification with large language models (LLMs) has recently attracted growing attention, due to their strong reasoning capabilities and transparent verification processes compared to traditional answer-only judgments. However, existing approaches to online claim verification, which requires iterative evidence retrieval and reasoning, still mainly rely on prompt engineering or pre-designed reasoning workflows, without unified training to improve necessary skills. Therefore, we introduce Veri-R1, an online reinforcement learning (RL) framework that enables an LLM to interact with a search engine and to receive reward signals that explicitly shape its planning, retrieval, and reasoning behaviors. This dynamic interaction of LLM with retrieval systems more accurately reflects real-world verification scenarios and fosters comprehensive verification skills. Empirical results show that Veri-R1 improves joint accuracy by up to 30% and doubles the evidence score, often surpassing its larger-scale model counterparts. Ablation studies further reveal the impact of reward components, and the link between output logits and label accuracy. Our results highlight the effectiveness of online RL for precise and faithful claim verification, providing an important foundation for future research. We release our code to support community progress in LLM empowered claim verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。