arXiv:2608.24350cs.CLcs.AI2026-08

让大模型推理更可信,通过精准分配事实验证的奖励

FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision

论文配图:FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision
图 1 · 摘自论文原文
  • 将事实验证粒度对齐到令牌级别,实现精准奖励定位
  • 用反事实证据分析判断事实信号可靠性,动态加权奖励
  • 在多个推理任务中显著提升事实正确率,不牺牲通用推理能力

为降低基于可验证奖励的强化学习训练中大型语言模型的幻觉风险,现有方法引入过程级事实监督。但由于事实信号粗粒度聚合及缺乏可靠性评估,导致事实验证与策略更新之间存在偏差。我们称此为噪声事实信用分配,并将其分解为信用定位模糊与信用可靠性模糊。为此,我们提出FARCA(事实对齐的可靠性感知信用分配)框架,将事实监督转化为局部化、可靠性加权的令牌级训练信号。FARCA通过将事实验证粒度与策略更新对齐,实现细粒度信用定位;并引入反事实证据归因,利用事实判断对关键证据的依赖性作为验证可靠性的经验代理,计算可靠性权重。这些权重调节事实奖励与局部策略优势,降低潜在不可靠信号对策略优化的影响。在不同模型和多个事实推理基准上的实验表明,FARCA显著提升模型事实性,同时保持通用推理能力。

原文摘要 · Abstract (English)

To reduce the hallucination risk caused by outcome-driven rewards in large language models trained through reinforcement learning with verifiable rewards, existing mitigation approaches introduce process-level factual supervision. However, due to coarse-grained aggregation of factual signals and the lack of reliability assessment for these signals, they create a mismatch between fact verification and policy updates. We term this noisy factual credit assignment and decompose it into two aspects: credit localization ambiguity and credit reliability ambiguity. To address these issues, we propose FARCA (Fact-Aligned Reliability-Aware Credit Assignment), a policy optimization framework that transforms factual supervision into localized, reliability-weighted token-level training signals. FARCA achieves fine-grained credit localization by aligning the granularity of fact verification with that of policy updates. It further introduces counterfactual evidence attribution, which uses the dependence of a factual judgment on key evidence as an empirical proxy for verification reliability to compute reliability weights. These weights modulate factual rewards and local policy advantages, reducing the influence of potentially unreliable signals on policy optimization. Experiments across different models and multiple factual reasoning benchmarks show that FARCA significantly improves model factuality while preserving general reasoning capabilities.

强化学习事实性信用分配语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。