用强化学习训练赛车策略,实现在真实车上的零样本部署并超越人类高手。
On learning racing policies with reinforcement learning
- 通过领域随机化与执行器建模,提升策略在真实环境的泛化能力。
- 在F1TENTH平台上,性能超过最优模型预测控制和人类专家驾驶。
- 为自动驾驶控制策略设计提供关键实践经验,适合关注鲁棒强化学习的开发者。
全自动驾驶车辆有望提升安全性和效率。然而,在复杂极端场景中确保可靠运行,需要能够逼近车辆极限的控制算法。本文以自动驾驶赛车任务为切入点,提出使用强化学习(RL)学习赛车策略。该方法结合领域随机化、执行器动力学建模和策略架构设计,实现策略在真实平台上的可靠、安全零样本部署。在F1TENTH赛车平台上的评估表明,所提出的RL策略不仅优于当前最先进的模型预测控制(MPC),而且据我们所知,首次在遥控赛车中实现超越人类专家驾驶员的表现。本工作揭示了推动性能提升的关键因素,为基于强化学习的自动驾驶控制策略设计提供了重要启示。
原文摘要 · Abstract (English)
Fully autonomous vehicles promise enhanced safety and efficiency. However, ensuring reliable operation in challenging corner cases requires control algorithms capable of performing at the vehicle limits. We address this requirement by considering the task of autonomous racing and propose solving it by learning a racing policy using Reinforcement Learning (RL). Our approach leverages domain randomization, actuator dynamics modeling, and policy architecture design to enable reliable and safe zero-shot deployment on a real platform. Evaluated on the F1TENTH race car, our RL policy not only surpasses a state-of-the-art Model Predictive Control (MPC), but, to the best of our knowledge, also represents the first instance of an RL policy outperforming expert human drivers in RC racing. This work identifies the key factors driving this performance improvement, providing critical insights for the design of robust RL-based control strategies for autonomous vehicles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。