arXiv:2608.10403cs.AI2026-08

让强化学习自动驾汽车主动发现并练习应对危险场景。

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

  • 根据当前驾驶策略弱点,智能生成针对性危险场景。
  • 在400万公里仿真测试中显著提升安全表现。
  • 适合需要高效安全训练的自动驾驶研发团队。

强化学习在自动驾驶中表现优异,但在线学习过程中保障策略安全仍具挑战,因难以充分接触危险罕见交通场景。真实交通环境具有长尾分布特性,传统采样难以覆盖高危交互,限制了强化学习策略对鲁棒安全行为的学习能力。现有方法通过合成挑战性或对抗性场景提升训练多样性,但通常将场景生成与策略演化分离,未显式建模生成扰动与当前策略弱点之间的关系。本文提出威胁引导的策略感知场景扰动(TPSP)方法,引入策略感知场景编码器,捕捉策略行为与周围环境的交互,实现与当前策略匹配的场景扰动。基于该表征,TPSP仅选择性扰动关键物体,而非全局均匀修改。进一步设计威胁引导优化策略,通过比较原始场景与扰动场景下策略轨迹的威胁等级差异,指导生成更具训练价值的安全关键场景。大量实验表明,TPSP在约400万公里的NAVSIM v2仿真数据上显著提升安全学习效率,消融实验验证:策略感知的定向扰动比随机或无策略感知策略提供更丰富的安全关键体验,可在有限交互预算下实现更安全的驾驶表现。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challenging scenes or adversarial situations. However, these approaches typically optimize scene generation objectives separately from the evolving policy, without explicitly modeling how generated perturbations relate to the current policy's weaknesses and learning needs. In this paper, we propose Threat-guided Policy-aware Scene Perturbation (TPSP) for safe autonomous driving with online RL. TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy. Based on this representation, TPSP selectively perturbs critical objects rather than applying uniform modifications across the scene. Furthermore, we develop a threat-guided optimization strategy that evaluates perturbed scenes through threat-level differences between policy rollouts on original and perturbed scenes, guiding the generation of safety-critical scenes with higher training value. Comprehensive experiments demonstrate that TPSP improves safety learning efficiency, achieving strong safety performance on NAVSIM v2 with approximately 4 million kilometers of simulated driving data. Ablation studies verify that policy-aware targeted perturbations provide more informative safety-critical experiences than random or policy-unaware strategies, enabling safer driving under limited interaction budgets.

自动驾驶强化学习安全训练场景生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。