arXiv:2606.05263cs.LGcs.AI2026-06

让语言智能体的决策过程可验证,防止作弊和错误推理。

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents

论文配图:Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents
图 1 · 摘自论文原文
  • 用反事实归因评估每步操作对最终成功的真实贡献。
  • 任务成功率提升至78.9%,证据准确率提高到82.8%。
  • 适合需要可靠推理与工具使用的长周期语言任务研究者。

基于可验证奖励的强化学习虽能提升推理与工具使用能力,但长周期语言智能体仍存在支持性证据链缺失、信念漂移和捷径行为等问题。现有过程奖励多为相关性设计,无法衡量某一步是否在特定干预下真正促成最终验证成功。本文提出CVT-RL,一种带密集可验证奖励、干预有效性门控和策略条件反事实贡献(PCCC)估计器的约束策略梯度算法。通过删除、语义替换、证据替换及工具输出扰动定义独立可控干预,从冻结参考策略采样延续路径,并采用选择调整的双重稳健估计器增强优势。信念控制仅依赖前缀可观测标签,而增强拉格朗日方法约束未支持声明、跳过验证、工具篡改与不安全调用。在长上下文问答、ALFWorld、ScienceWorld及网页/工具任务上,CVT-RL将平均任务成功率从计算匹配的非因果强化学习的71.8%、信息匹配反事实基线的75.4%提升至78.9%,证据F1从78.9%升至82.8%,测量到的漏洞利用从7.2%降至3.9%。独立人工审计显示CVT-RL漏洞率为4.6%,低于基线的8.1%,自适应探测逃避攻击也仅使漏洞率升至7.1%。分层自助法与混合效应检验在霍尔姆校正后所有主指标p<0.01。精心设计的反事实信用机制结合有效性门控、诊断与可验证约束,为长周期语言智能体强化学习提供了可复现的可靠性路径。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chains, belief drift, and shortcut actions that satisfy terminal checks. Existing process rewards are mostly correlational: they reward retrieval-, reflection-, or verification-like steps without estimating whether the step contributes to final verified success under a specified intervention. We propose CVT-RL, a constrained policy-gradient algorithm with dense verifiable rewards, intervention-validity gating, and a policy-conditioned counterfactual contribution (PCCC) estimator. Deletion, semantic substitution, evidence substitution, and tool-output perturbation define separate controlled interventions; continuations are sampled from a frozen reference policy, and a selection-adjusted doubly robust estimator augments the advantage. Belief control uses only prefix-observable labels, while an augmented Lagrangian constrains unsupported claims, skipped verification, tool tampering, and unsafe calls. On long-context QA, ALFWorld, ScienceWorld, and web/tool tasks, CVT-RL improves average task success from 71.8% for compute-matched non-causal RL and 75.4% for an information-matched counterfactual-process baseline to 78.9%, improves evidence F1 from 78.9 to 82.8 over the information-matched baseline, and reduces measured hacking from 7.2% to 3.9%. Independent human audit estimates 4.6% hacking for CVT-RL versus 8.1% for the information-matched baseline, and adaptive detector-evasion attacks raise hacking only to 7.1%. Stratified bootstrap and mixed-effects tests give p<0.01 after Holm correction for all primary metrics. Carefully scoped counterfactual credit, paired with validity gating, diagnostics, and verifiable constraints, provides a reproducible route toward more reliable long-horizon RL for language agents.

强化学习可验证性语言代理反事实分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。