arXiv:2503.15200cs.LG2025-03ICML被引 12

用记忆痕迹提升部分可观测强化学习的样本效率。

Partially Observable Reinforcement Learning with Memory Traces

  • 用指数移动平均构建历史观测的紧凑表示
  • 理论证明记忆痕迹可降低回报误差,提升学习效率
  • 适合长期依赖环境下的在线强化学习任务

部分可观测环境在强化学习中带来巨大计算挑战,需考虑长时间历史。有限观察窗口的学习随窗口增长迅速变得不可行。本文引入记忆痕迹:受适用性痕迹启发,以指数移动平均形式表示历史观测的紧凑表征。我们为离线策略评估问题建立了样本复杂度边界,量化了具有李普希茨连续价值估计时,记忆痕迹实现的回报误差。揭示其与窗口方法的紧密联系,并在特定环境中证明记忆痕迹显著更高效。最后,通过在线强化学习实验验证了记忆痕迹在价值预测和控制任务中的有效性。

原文摘要 · Abstract (English)

Partially observable environments present a considerable computational challenge in reinforcement learning due to the need to consider long histories. Learning with a finite window of observations quickly becomes intractable as the window length grows. In this work, we introduce memory traces. Inspired by eligibility traces, these are compact representations of the history of observations in the form of exponential moving averages. We prove sample complexity bounds for the problem of offline on-policy evaluation that quantify the return errors achieved with memory traces for the class of Lipschitz continuous value estimates. We establish a close connection to the window approach, and demonstrate that, in certain environments, learning with memory traces is significantly more sample efficient. Finally, we underline the effectiveness of memory traces empirically in online reinforcement learning experiments for both value prediction and control.

强化学习记忆机制部分可观测在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。