arXiv:2503.17409cs.LGcs.RO2025-03中稿 · AAAI

用高斯过程建模奖励,让稀疏奖励变密集有效

Reward Redistribution via Gaussian Process Likelihood Estimation

  • 用高斯过程捕捉状态动作对间的依赖关系
  • 通过留一法最大化轨迹回报似然,提升样本效率
  • 适合长周期稀疏奖励的强化学习任务

在许多实际强化学习任务中,反馈仅在长时程结束时提供,导致奖励稀疏且延迟。现有奖励重分配方法通常假设每步奖励相互独立,忽略了状态动作对之间的关联性。本文提出基于高斯过程的似然奖励重分配(GP-LRR)框架,将奖励函数视为高斯过程的一个样本,通过核函数显式建模状态动作对间的依赖关系。通过采用留一法策略,最大化观测到的轨迹回报似然,该框架天然引入不确定性正则化。此外,我们证明当使用退化核且无观测噪声时,传统基于均方误差(MSE)的奖励重分配可视为本框架的特例。将GP-LRR与如软演员-批评家(Soft Actor-Critic)等离策略算法结合,在多个MuJoCo基准测试中,生成密集且信息丰富的奖励信号,显著提升样本效率与策略性能。

原文摘要 · Abstract (English)

In many practical reinforcement learning tasks, feedback is only provided at the end of a long horizon, leading to sparse and delayed rewards. Existing reward redistribution methods typically assume that per-step rewards are independent, thus overlooking interdependencies among state-action pairs. In this paper, we propose a Gaussian process based Likelihood Reward Redistribution (GP-LRR) framework that addresses this issue by modeling the reward function as a sample from a Gaussian process, which explicitly captures dependencies between state-action pairs through the kernel function. By maximizing the likelihood of the observed episodic return via a leave-one-out strategy that leverages the entire trajectory, our framework inherently introduces uncertainty regularization. Moreover, we show that conventional mean-squared-error (MSE) based reward redistribution arises as a special case of our GP-LRR framework when using a degenerate kernel without observation noise. When integrated with an off-policy algorithm such as Soft Actor-Critic, GP-LRR yields dense and informative reward signals, resulting in superior sample efficiency and policy performance on several MuJoCo benchmarks.

强化学习奖励重分配高斯过程稀疏奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。