arXiv:2504.08161cs.LGcs.AI2025-04被引 19

挑战强化学习传统范式,为持续学习提出新框架

Rethinking the Foundations for Continual Reinforcement Learning

  • 用历史过程替代马尔可夫决策过程,支持持续学习
  • 改用偏差遗憾作为评估指标,更适配长期学习目标
  • 适合研究持续学习与智能体长期适应性的学者

传统强化学习的目标是寻找最大化期望奖励总和的最优策略,一旦找到即停止学习。这与持续强化学习(continual reinforcement learning)相悖,后者要求智能体持续学习与适应。尽管二者差异明显,当前持续学习进展仍受传统范式影响。本文分析发现,传统强化学习的四个核心基础——马尔可夫决策过程形式化、关注非时间性成果、以期望奖励总和为评估标准、以及采用嵌套这些基础的分段基准环境——均与持续学习目标相冲突。为此,论文提出新形式化框架,摒弃前两者,以历史过程作为数学形式化,并引入专为持续学习设计的偏差遗憾评估指标。最后讨论如何进一步摆脱其余两个基础的影响。

原文摘要 · Abstract (English)

In the traditional view of reinforcement learning, the agent's goal is to find an optimal policy that maximizes its expected sum of rewards. Once the agent finds this policy, the learning ends. This view contrasts with \emph{continual reinforcement learning}, where learning does not end, and agents are expected to continually learn and adapt indefinitely. Despite the clear distinction between these two paradigms of learning, much of the progress in continual reinforcement learning has been shaped by foundations rooted in the traditional view of reinforcement learning. In this paper, we first examine whether the foundations of traditional reinforcement learning are suitable for the continual reinforcement learning paradigm. We identify four key pillars of the traditional reinforcement learning foundations that are antithetical to the goals of continual learning: the Markov decision process formalism, the focus on atemporal artifacts, the expected sum of rewards as an evaluation metric, and episodic benchmark environments that embrace the other three foundations. We then propose a new formalism that sheds the first and the third foundations and replaces them with the history process as a mathematical formalism and a new definition of deviation regret, adapted for continual learning, as an evaluation metric. Finally, we discuss possible approaches to shed the other two foundations.

强化学习持续学习形式化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。