在自动驾驶泊车场景中揭示持续强化学习的四大挑战
Continual Reinforcement Learning for Cyber-Physical Systems: Lessons Learned and Open Challenges
- 用PPO算法逐个训练四种角度泊车任务,模拟持续学习环境
- 发现灾难性遗忘、超参敏感、抽象能力不足等关键问题
- 适合研究持续学习与神经网络鲁棒性的学者参考
持续学习(CL)旨在使智能体能够适应并泛化已有技能,以应对新任务或环境变化,尤其适用于动态多任务系统。本文通过在自动驾驶泊车环境中进行实验,揭示了持续强化学习(CRL)中的若干开放挑战。实验设置为依次训练四个不同角度的泊车场景,使用近端策略优化(PPO)算法。结果表明:环境抽象能力不足、超参数高度敏感、灾难性遗忘现象显著,以及神经网络容量利用效率低。这些发现凸显了当前神经网络在持续学习中的局限性,并提出需解决的关键研究问题。此外,研究呼吁跨学科合作,尤其是计算机科学与神经科学的融合,以推动鲁棒性持续学习系统的发展。
原文摘要 · Abstract (English)
Continual learning (CL) is a branch of machine learning that aims to enable agents to adapt and generalise previously learned abilities so that these can be reapplied to new tasks or environments. This is particularly useful in multi-task settings or in non-stationary environments, where the dynamics can change over time. This is particularly relevant in cyber-physical systems such as autonomous driving. However, despite recent advances in CL, successfully applying it to reinforcement learning (RL) is still an open problem. This paper highlights open challenges in continual RL (CRL) based on experiments in an autonomous driving environment. In this environment, the agent must learn to successfully park in four different scenarios corresponding to parking spaces oriented at varying angles. The agent is successively trained in these four scenarios one after another, representing a CL environment, using Proximal Policy Optimisation (PPO). These experiments exposed a number of open challenges in CRL: finding suitable abstractions of the environment, oversensitivity to hyperparameters, catastrophic forgetting, and efficient use of neural network capacity. Based on these identified challenges, we present open research questions that are important to be addressed for creating robust CRL systems. In addition, the identified challenges call into question the suitability of neural networks for CL. We also identify the need for interdisciplinary research, in particular between computer science and neuroscience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。