用强化学习统一优化事实核查的多阶段流程,提升准确率与效率。
From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification

- 设计统一策略,动态协调分解、查证、生成与判断各阶段
- 引入过程感知奖励,解决最终标签反馈稀疏问题,提升各阶段学习效果
- 在多个数据集上超越基线模型,适合需要高可靠性的事实核查场景
近期结合大语言模型与检索增强推理的方法在自动化事实核查中展现出潜力。为处理复杂声明,现有核查流程通常采用多阶段工作流,协调紧密耦合的模块,包括声明分解、证据获取和结论预测。然而,现有方法通常孤立优化各阶段或依赖固定启发式规则,限制了阶段间的自适应协调,可能导致次优结果。本文提出 ProFact,一种面向多阶段事实核查轨迹的智能体强化学习框架,实现端到端优化。ProFact 训练统一策略以协调声明分解、证据搜索、答案生成和结论预测。针对最终真实性标签提供的稀疏且延迟的监督信号,ProFact 引入过程感知奖励,在整个核查过程中提供阶段级学习信号。实证评估表明,ProFact 在验证性能和推理效率上均持续优于强基线模型。结果表明,过程感知轨迹优化对多阶段事实核查具有显著有效性。
原文摘要 · Abstract (English)
Recent approaches combining Large Language Models (LLMs) with retrieval-augmented reasoning have shown promise for automated fact verification. To process complex claims, these verification pipelines typically execute multi-stage workflows that coordinate tightly coupled modules, including claim decomposition, evidence gathering, and verdict prediction. However, existing methods optimize individual stages in isolation or rely on fixed heuristics, which limits adaptive coordination among stages and can lead to suboptimal outcomes. In this work, we propose ProFact, an agentic reinforcement learning framework for end-to-end optimization of multi-stage fact verification trajectories. ProFact trains a unified policy to coordinate claim decomposition, evidence seeking, answer generation, and verdict prediction. To address the sparse and delayed supervision provided by final veracity labels, ProFact introduces process-aware rewards that provide stage-level learning signals throughout the verification process. Empirical evaluation shows that ProFact consistently outperforms strong baselines in both verification performance and inference efficiency. These results highlight the effectiveness of process-aware trajectory optimization for multi-stage fact verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。