arXiv:2505.18347cs.LGcs.AI2025-05被引 2

用游戏Agar.io构建持续强化学习新基准,挑战智能体长期适应能力。

The Cell Must Go On: Agar.io for Continual Reinforcement Learning

  • 基于非回合制的Agar.io设计持续学习平台AgarCL,具高维动态环境。
  • 三种主流强化学习算法在主任务上表现有限,说明挑战超出传统稳定性难题。
  • 拆解小任务分析环境各成分影响,为持续学习研究提供精细评估工具。

持续强化学习(Continual RL)要求智能体持续学习而非收敛后固定策略。现有方法多通过修改回合制环境或设计特定模拟器来模拟变化,但主要捕捉突发性数据流变化,且仍依赖回合结构。少数专门设计的模拟器又规模有限。本文提出AgarCL,基于游戏Agar.io构建持续学习平台:其为非回合制、高维、随机演化、连续动作与部分可观测环境。我们提供DQN、PPO、SAC在主任务上的基准结果,并在多个子任务中评估,以分离不同环境组件带来的挑战。进一步测试了三种持续学习方法(Shrink and Perturb、ReDo、Continual Backpropagation),发现其性能提升有限,表明AgarCL的挑战已超越稳定性-可塑性困境。

原文摘要 · Abstract (English)

Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fixed for evaluation. This setting is well-suited to environments that the agent perceives as changing over time, rendering any static policy ineffective. In continual RL, researchers often simulate such changes either by modifying episodic environments to incorporate task shifts during interaction or by designing simulators that explicitly model continual dynamics. However, transforming episodic problems into continual ones primarily captures scenarios involving abrupt changes in the data stream and still relies on episodic structure. Meanwhile, the few simulators explicitly designed for empirical continual RL research are often limited in scope or complexity. In this paper, we introduce AgarCL, a research platform for continual RL that enables agents to progress toward increasingly sophisticated behaviour. AgarCL is based on the game Agar.io, a non-episodic, high-dimensional problem with stochastic, ever-evolving dynamics, continuous actions, and partial observability. We provide benchmark results for DQN, PPO, and SAC on the primary continual RL challenge, as well as across a suite of smaller tasks within AgarCL. These smaller tasks isolate aspects of the full environment and allow us to characterize the distinct challenges posed by different components of the game. We further evaluate three continual learning methods-Shrink and Perturb, ReDo, and Continual Backpropagation-and observe little improvement over standard RL algorithms, suggesting that the challenges posed by AgarCL extend beyond the stability-plasticity dilemma.

持续学习强化学习游戏基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。