用强化学习生成挑战性交通场景,自动测试并提升自动驾驶安全性。
CRASH: Challenging Reinforcement-Learning Based Adversarial Scenarios For Safety Hardening
- 用对抗性强化学习控制虚拟车辆,自动制造碰撞场景
- 使规则与学习型规划器碰撞率超90%,安全加固后碰撞率降26%
- 适合自动驾驶安全测试与算法鲁棒性评估的研究者
确保自动驾驶汽车(AV)的安全性需要识别出道路测试无法发现的罕见但关键的故障案例。高保真仿真提供了可扩展的替代方案,但自动生成能有效压力测试AV运动规划器的逼真且多样化的交通场景仍是核心挑战。本文提出CRASH——一种基于对抗性深度强化学习的框架,用于生成挑战性场景以强化安全。首先,CRASH可在AV仿真环境中控制非玩家角色(NPC)代理,自动引发与自车的碰撞,从而暴露其运动规划器的缺陷。我们还提出一种新颖的“安全加固”方法,通过在对抗性代理的威胁下反复迭代优化规划器,利用失败案例增强整车系统。实验在简化双车道高速公路场景中进行,结果表明该方法能以超过90%的碰撞率破坏规则与学习型规划器;经安全加固后,自车碰撞率降低26%。初步结果表明,基于强化学习的安全加固是一种有前景的、面向场景驱动的仿真测试方法。
原文摘要 · Abstract (English)
Ensuring the safety of autonomous vehicles (AVs) requires identifying rare but critical failure cases that on-road testing alone cannot discover. High-fidelity simulations provide a scalable alternative, but automatically generating realistic and diverse traffic scenarios that can effectively stress test AV motion planners remains a key challenge. This paper introduces CRASH - Challenging Reinforcement-learning based Adversarial scenarios for Safety Hardening - an adversarial deep reinforcement learning framework to address this issue. First CRASH can control adversarial Non Player Character (NPC) agents in an AV simulator to automatically induce collisions with the Ego vehicle, falsifying its motion planner. We also propose a novel approach, that we term safety hardening, which iteratively refines the motion planner by simulating improvement scenarios against adversarial agents, leveraging the failure cases to strengthen the AV stack. CRASH is evaluated on a simplified two-lane highway scenario, demonstrating its ability to falsify both rule-based and learning-based planners with collision rates exceeding 90%. Additionally, safety hardening reduces the Ego vehicle's collision rate by 26%. While preliminary, these results highlight RL-based safety hardening as a promising approach for scenario-driven simulation testing for autonomous vehicles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。