arXiv:2602.12342cs.LGcs.AI2026-02被引 6

用语言模型的信念变化奖励中间进展,提升长时序任务表现。

Intrinsic Credit Assignment for Long Horizon Interaction

  • 通过模型对目标解概率的变化衡量信用分配
  • 在长时交互中性能随测试时交互增加而持续提升
  • 适合需要探索与长期规划的智能体训练场景

如何让智能体在长时序中应对不确定性?本文提出ΔBelief-RL,利用语言模型自身对目标解的概率变化作为内在奖励信号,以评估中间进展。通过合成交互数据训练,该方法使智能体具备信息探索能力,在客户客服、个性化推荐等分布外任务中均优于纯结果导向奖励。值得注意的是,随着测试时交互长度超过训练范围,性能仍持续提升,且在Pass@k指标上表现出更高的交互效率。本工作提供了一种可扩展的长时序不确定性应对训练策略,通过内在ΔBelief奖励实现对中间动作的有效信用分配。

原文摘要 · Abstract (English)

How can we train agents to navigate uncertainty over long horizons? In this work, we propose ΔBelief-RL, which leverages a language model's own intrinsic beliefs to reward intermediate progress. Our method utilizes the change in the probability an agent assigns to the target solution for credit assignment. By training on synthetic interaction data, ΔBelief-RL teaches information-seeking capabilities that consistently outperform purely outcome-based rewards for Reinforcement Learning, with improvements generalizing to out-of-distribution applications ranging from customer service to personalization. Notably, the performance continues to improve as we scale test-time interactions beyond the training horizon, with interaction-efficiency increasing even on Pass@k metrics. Overall, our work introduces a scalable training strategy for navigating uncertainty over a long-horizon, by enabling credit assignment to intermediate actions via intrinsic ΔBelief rewards.

强化学习长时序决策信用分配语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。