arXiv:2601.01800cs.LGcs.AI2026-01

针对自动驾驶中稀疏高危风险,提出更聚焦的对抗训练方法,提升安全性。

Sparse Threats, Focused Defense: Criticality-Aware Robust Reinforcement Learning for Safe Autonomous Driving

  • 设计双组件对抗框架,区分风险暴露与鲁棒学习。
  • 碰撞率降低至少22.66%,优于现有最先进方法。
  • 适合关注自动驾驶安全性的研究者与工程团队。

强化学习在自动驾驶中潜力巨大,但对扰动的脆弱性仍是实际部署的关键障碍。现有对抗训练多将交互建模为连续攻击的零和博弈,忽略了智能体与攻击方的不对称性,且未反映安全关键风险的稀疏性,导致鲁棒性不足。为此,本文提出关键性感知鲁棒强化学习(CARRL),包含风险暴露代理(REA)与风险靶向鲁棒智能体(RTRA)。二者构成非零和博弈:REA通过解耦优化机制,在有限预算下聚焦探测稀疏的安全关键时刻(如碰撞);RTRA则利用双回放缓冲区融合良性与对抗经验,并通过策略一致性约束稳定行为。实验表明,该方法在所有测试场景中碰撞率至少降低22.66%,显著优于当前最优基线。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has shown considerable potential in autonomous driving (AD), yet its vulnerability to perturbations remains a critical barrier to real-world deployment. As a primary countermeasure, adversarial training improves policy robustness by training the AD agent in the presence of an adversary that deliberately introduces perturbations. Existing approaches typically model the interaction as a zero-sum game with continuous attacks. However, such designs overlook the inherent asymmetry between the agent and the adversary and then fail to reflect the sparsity of safety-critical risks, rendering the achieved robustness inadequate for practical AD scenarios. To address these limitations, we introduce criticality-aware robust RL (CARRL), a novel adversarial training approach for handling sparse, safety-critical risks in autonomous driving. CARRL consists of two interacting components: a risk exposure adversary (REA) and a risk-targeted robust agent (RTRA). We model the interaction between the REA and RTRA as a general-sum game, allowing the REA to focus on exposing safety-critical failures (e.g., collisions) while the RTRA learns to balance safety with driving efficiency. The REA employs a decoupled optimization mechanism to better identify and exploit sparse safety-critical moments under a constrained budget. However, such focused attacks inevitably result in a scarcity of adversarial data. The RTRA copes with this scarcity by jointly leveraging benign and adversarial experiences via a dual replay buffer and enforces policy consistency under perturbations to stabilize behavior. Experimental results demonstrate that our approach reduces the collision rate by at least 22.66\% across all cases compared to state-of-the-art baseline methods.

自动驾驶强化学习对抗训练安全关键

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。