arXiv:2601.10306cs.AIcs.CL2026-01ACL被引 7

通过强化学习提升长文本推理中的证据检索质量,避免盲目猜测。

Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning

  • 设计新算法EAPO,用分组相对证据奖励实现细粒度过程监督。
  • 在8个基准上显著优于当前最佳模型,证据检索准确率明显提升。
  • 适合需要高可靠性推理的场景,如金融、医疗问答系统。

尽管强化学习已推动大模型推理发展,但在长上下文场景中,结果奖励稀疏的问题导致无法惩罚无依据的‘幸运猜测’,使得关键的‘大海捞针’式证据检索缺乏有效监督。为此,本文提出EAPO(证据增强策略优化)方法。首先建立证据增强推理范式,通过树状证据采样验证:精确的证据提取是长上下文推理的决定性瓶颈。基于此,EAPO引入专用强化学习算法,由奖励模型计算分组相对证据奖励,提供密集的过程监督以显式提升证据质量。为维持训练全程的精准监督,进一步引入自适应奖励-策略共进化机制,通过结果一致的回放迭代优化奖励模型,增强其判别能力以确保精确的过程引导。在八个基准上的综合评估表明,相比现有最优基线,EAPO显著提升了长上下文推理性能。

原文摘要 · Abstract (English)

While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long-context scenarios is hindered by sparsity of outcome rewards. This limitation fails to penalize ungrounded "lucky guesses," leaving the critical process of needle-in-a-haystack evidence retrieval largely unsupervised. To address this, we propose EAPO (Evidence-Augmented Policy Optimization). We first establish the Evidence-Augmented Reasoning paradigm, validating via Tree-Structured Evidence Sampling that precise evidence extraction is the decisive bottleneck for long-context reasoning. Guided by this insight, EAPO introduces a specialized RL algorithm where a reward model computes a Group-Relative Evidence Reward, providing dense process supervision to explicitly improve evidence quality. To sustain accurate supervision throughout training, we further incorporate an Adaptive Reward-Policy Co-Evolution mechanism. This mechanism iteratively refines the reward model using outcome-consistent rollouts, sharpening its discriminative capability to ensure precise process guidance. Comprehensive evaluations across eight benchmarks demonstrate that EAPO significantly enhances long-context reasoning performance compared to SOTA baselines.

强化学习长文本推理证据提取大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。