提出可预测环境变化导致性能下降的评估指标,提升持续学习强化学习效果。
CHIRPs: Change-Induced Regret Proxy metrics for Lifelong Reinforcement Learning
- 基于环境变化设计后悔代理指标,量化变化对智能体影响。
- 在基准测试中性能比次优方法高48%,10项任务中有8项达成最佳成功率。
- 适用于需要应对动态环境的持续学习场景,如机器人与自适应系统。
强化学习智能体训练成本高且对环境变化敏感。当任务频繁变动时,其表现常显著下降,限制了在真实世界中的广泛应用。尽管已有多种持续强化学习方法缓解灾难性遗忘或实现正向迁移,但尚无研究能从变化本身预测对智能体性能的影响。理解这一关系有助于智能体主动应对变化,提升学习效率。本文提出改变引发的后悔代理(CHIRP)指标,建立变化与性能下降之间的联系,并在两个环境中验证其有效性。基于CHIRP的简单智能体在首个基准测试中性能比最优对比方法高出48%,在第二个复杂基准测试中,10项任务中有8项取得最高成功率,证明其对现有持续学习方法的显著改进。
原文摘要 · Abstract (English)
Reinforcement learning (RL) agents are costly to train and fragile to environmental changes. They often perform poorly when there are many changing tasks, prohibiting their widespread deployment in the real world. Many Lifelong RL agent designs have been proposed to mitigate issues such as catastrophic forgetting or demonstrate positive characteristics like forward transfer when change occurs. However, no prior work has established whether the impact on agent performance can be predicted from the change itself. Understanding this relationship will help agents proactively mitigate a change's impact for improved learning performance. We propose Change-Induced Regret Proxy (CHIRP) metrics to link change to agent performance drops and use two environments to demonstrate a CHIRP's utility in lifelong learning. A simple CHIRP-based agent achieved $48\%$ higher performance than the next best method in one benchmark and attained the best success rates in 8 of 10 tasks in a second benchmark which proved difficult for existing lifelong RL agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。