用强化学习微调行为模型,提升自动驾驶仿真中代理的可靠性。
Improving Agent Behaviors with RL Fine-tuning for Autonomous Driving
- 通过闭环强化学习微调行为模型,缓解部署时的分布偏移问题。
- 在Waymo开放仿真代理挑战中,碰撞率等关键指标显著降低。
- 新设计评估基准,可直接检验仿真代理对自动驾驶规划器的评价能力。
自主车辆研究中的主要挑战之一是建模交通参与者的行为,这对构建真实可靠的仿真环境(用于离线评估)和实现车载规划中的交通参与者轨迹预测至关重要。尽管监督学习已在多个领域成功建模了参与者行为,但这些模型在测试时可能因分布偏移而性能下降。本文提出通过强化学习进行闭环微调行为模型,以提升其可靠性。实验表明,该方法在Waymo Open Sim Agents挑战中不仅整体性能更优,且碰撞率等目标指标显著改善。此外,我们提出了一个新颖的策略评估基准,可直接衡量仿真代理对自动驾驶规划器质量的评估能力,并验证了本方法的有效性。
原文摘要 · Abstract (English)
A major challenge in autonomous vehicle research is modeling agent behaviors, which has critical applications including constructing realistic and reliable simulations for off-board evaluation and forecasting traffic agents motion for onboard planning. While supervised learning has shown success in modeling agents across various domains, these models can suffer from distribution shift when deployed at test-time. In this work, we improve the reliability of agent behaviors by closed-loop fine-tuning of behavior models with reinforcement learning. Our method demonstrates improved overall performance, as well as improved targeted metrics such as collision rate, on the Waymo Open Sim Agents challenge. Additionally, we present a novel policy evaluation benchmark to directly assess the ability of simulated agents to measure the quality of autonomous vehicle planners and demonstrate the effectiveness of our approach on this new benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。