arXiv:2606.04029cs.LGcs.AI2026-06中稿 · ICML

部署的强化学习应持续优化,而非训练后固定不动。

Position: Deployed Reinforcement Learning should be Continual

论文配图:Position: Deployed Reinforcement Learning should be Continual
图 1 · 摘自论文原文
  • 提出部署后需持续学习,打破训练即结束的旧模式。
  • 识别部署后四大非平稳因素,证明持续学习必要性。
  • 适合关注真实场景智能体长期适应的研究者与工程师。

强化学习在实际应用中日益受到关注和采用。目前多数系统遵循训练后固定的范式,即训练完成的智能体在与环境交互时不再学习,直到性能下降才重新训练。本文认为,部署一个未达最优但能接收评价奖励信号的智能体,本质上是一个持续强化学习问题。我们识别了部署后存在的四种非平稳性来源,表明必须进行永不终止的学习。通过分析现实世界中成功的持续强化学习案例,本文提出了摆脱当前训练后固定范式的优点与具体措施。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has received increasing attention and adoption in real-world use cases. Most of these systems follow a train-then-fix paradigm, where trained agents do not learn while interacting with the world until performance degrades and retraining becomes necessary. In this position paper, we argue that deploying an agent that is incapable of optimality, but receives an evaluative reward signal, is inherently a continual RL problem. We identify four sources of non-stationarity after deployment that necessitate never-ending learning, and highlight why the best deployed agents never stop adapting. We analyze successful examples of continual RL in the real world, and present the community with the advantages and measures to move away from the current train-then-fix paradigm.

强化学习持续学习部署系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。