轻量级环境Forager挑战持续强化学习中的部分可观测性问题
Forager: a lightweight testbed for continual learning with partial observability in RL

- 构建轻量级部分可观测环境,内存占用恒定
- 现有方法易失塑性,状态构建更有效
- 适合研究长期学习与记忆机制的算法
在持续强化学习(CRL)中,良好性能需要在大型部分可观测世界中持续学习、决策与探索。当前多数实验聚焦于‘可塑性丧失’问题,但在经典完全可观测马尔可夫决策过程(MDP)中加入非平稳性的一次性实验,忽略了部分可观测性和记忆/递归机制的重要性。原因之一是许多部分可观测的CRL环境计算成本过高。本文提出Forager,一个轻量级的部分可观测CRL环境,具有恒定内存开销。我们提供一系列实验和样本任务,证明Forager对现有CRL智能体具有挑战性,同时支持对其行为的深入分析。结果显示,智能体确实存在可塑性丧失,现有缓解策略有一定帮助,但最有效的是利用状态构造。最后,我们提出Forager的一个变体,能生成无限新任务流,清晰揭示当前CRL方法的局限性。
原文摘要 · Abstract (English)
In continual reinforcement learning (CRL), good performance requires never-ending learning, acting, and exploration in a big, partially observable world. Most CRL experiments have focused on loss of plasticity -- the inability to keep learning -- in one-off experiments where some unobservable non-stationarity is added to classic fully observable MDPs. Further, these experiments rarely consider the role of partial observability and the importance of CRL agents that use memory or recurrence. One potential reason for this focus on mitigating loss of plasticity without considering partial observability is that many partially-observable CRL environments are prohibitively expensive. In this paper, we introduce Forager, a light-weight partially-observable CRL environment with a constant memory footprint. We provide a set of experiments and sample tasks demonstrating that Forager is challenging for current CRL agents and yet also allows for in-depth study of those agents. We demonstrate that agents exhibit loss of plasticity, proposed mitigations can help, but that most useful is to leverage state construction. We conclude with a variant of Forager that generates an unending stream of new tasks to learn that clearly highlights the limitations of current CRL agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。