用物理信息强化学习,让汽车在稀疏奖励下自动安全漂移。
Autonomous Drifting Based on Maximal Safety Probability Learning
- 基于物理损失替代人工设计奖励,实现安全概率最大化
- 仅需稀疏二值奖励即可学会高风险场景下的安全行为
- 无需参考轨迹或复杂调奖,适合高速竞速等极限场景
本文提出一种基于最大安全概率学习的新型自动驾驶学习框架。高效学习依赖于能区分理想/非理想状态的有益奖励,但手动设计此类奖励因难以在众多安全状态中辨别优劣而极具挑战。另一方面,最大化安全概率的策略学习虽无需繁琐奖励设计,却因依赖稀疏时间上的二值奖励而数值上困难。本文表明,物理信息强化学习可高效学习此类最大安全策略。与现有漂移控制方法不同,本方法无需特定参考轨迹或复杂奖励塑造,仅通过稀疏二值奖励即可学习安全行为。这得益于物理损失项所起的类奖励塑造作用。在常规弯道保持车道和高速竞速场景下的安全漂移任务中验证了该方法的有效性。
原文摘要 · Abstract (English)
This paper proposes a novel learning-based framework for autonomous driving based on the concept of maximal safety probability. Efficient learning requires rewards that are informative of desirable/undesirable states, but such rewards are challenging to design manually due to the difficulty of differentiating better states among many safe states. On the other hand, learning policies that maximize safety probability does not require laborious reward shaping but is numerically challenging because the algorithms must optimize policies based on binary rewards sparse in time. Here, we show that physics-informed reinforcement learning can efficiently learn this form of maximally safe policy. Unlike existing drift control methods, our approach does not require a specific reference trajectory or complex reward shaping, and can learn safe behaviors only from sparse binary rewards. This is enabled by the use of the physics loss that plays an analogous role to reward shaping. The effectiveness of the proposed approach is demonstrated through lane keeping in a normal cornering scenario and safe drifting in a high-speed racing scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。