arXiv:2410.22752cs.ROcs.AI2024-10被引 1

用软约束提升自动驾驶强化学习的鲁棒性,减少模仿学习的保守行为。

SoftCTRL: Soft conservative KL-control of Transformer Reinforcement Learning for Autonomous Driving

  • 通过隐式熵-KL控制,动态调节模仿与强化学习的平衡。
  • 在未见城市场景中失败率降低17%以上,驾驶行为更接近人类。
  • 适合需要高安全性和自然驾驶行为的自动驾驶系统研发。

近年来,城市自动驾驶汽车(SDV)的运动规划因道路要素复杂交互而成为热门问题。许多方法依赖大规模人工采样数据,通过模仿学习(IL)进行处理。尽管有效,但仅靠模仿学习难以解决安全与可靠性问题。将模仿学习与强化学习(RL)结合,在RL损失中加入模仿与强化策略间的KL散度可缓解模仿学习的缺陷,但会因模仿学习的协变量偏移导致过度保守。为此,我们提出一种新方法,通过隐式熵-KL控制,以简单方式减轻过度保守特性。我们在多个未见过的模拟城市场景中验证了该方法,结果显示:虽然模仿学习在模仿任务中表现良好,但所提方法显著提升了鲁棒性(失败率降低超过17%),并生成类人驾驶行为。

原文摘要 · Abstract (English)

In recent years, motion planning for urban self-driving cars (SDV) has become a popular problem due to its complex interaction of road components. To tackle this, many methods have relied on large-scale, human-sampled data processed through Imitation learning (IL). Although effective, IL alone cannot adequately handle safety and reliability concerns. Combining IL with Reinforcement learning (RL) by adding KL divergence between RL and IL policy to the RL loss can alleviate IL's weakness but suffer from over-conservation caused by covariate shift of IL. To address this limitation, we introduce a method that combines IL with RL using an implicit entropy-KL control that offers a simple way to reduce the over-conservation characteristic. In particular, we validate different challenging simulated urban scenarios from the unseen dataset, indicating that although IL can perform well in imitation tasks, our proposed method significantly improves robustness (over 17\% reduction in failures) and generates human-like driving behavior.

自动驾驶强化学习模仿学习轨迹规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。