用专家指导的自适应攻击法,更高效稳定地破解自动驾驶强化学习模型
Sharpening the Spear: Adaptive Expert-Guided Adversarial Attack Against DRL-based Autonomous Driving Policies
- 通过模仿成功攻击样本生成专家策略,提升攻击方向性
- 在复杂场景下碰撞率更高,训练更稳定,即使专家不完美也有效
- 适合研究自动驾驶安全漏洞或对抗样本防御的工程师与研究员
深度强化学习(DRL)已成为自动驾驶的前沿范式,但其仍极易受对抗攻击影响,威胁实际部署安全。现有攻击方法常依赖高频攻击,而真实攻击机会多为情境相关且时间稀疏,导致效率低下;降低频率虽能提升效率,却因对手探索不足造成训练不稳定。为此,本文提出一种自适应专家引导的对抗攻击方法,兼顾攻击稳定性与效率。首先通过模仿学习从成功攻击示范中提取专家策略,并采用集成的Mixture-of-Experts架构增强跨场景泛化能力;随后利用KL散度正则化项引导基于DRL的攻击者。针对专家策略可能不完善的挑战,引入性能感知退火策略,随攻击者表现提升逐步减少对专家的依赖。大量实验表明,该方法在碰撞率、攻击效率和训练稳定性上均优于现有方法,尤其在专家策略不佳时仍具鲁棒性。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) has emerged as a promising paradigm for autonomous driving. However, despite their advanced capabilities, DRL-based policies remain highly vulnerable to adversarial attacks, posing serious safety risks in real-world deployments. Investigating such attacks is crucial for revealing policy vulnerabilities and guiding the development of more robust autonomous systems. While prior attack methods have made notable progress, they still face several challenges: 1) they often rely on high-frequency attacks, yet critical attack opportunities are typically context-dependent and temporally sparse, resulting in inefficient attack patterns; 2) restricting attack frequency can improve efficiency but often results in unstable training due to the adversary's limited exploration. To address these challenges, we propose an adaptive expert-guided adversarial attack method that enhances both the stability and efficiency of attack policy training. Our method first derives an expert policy from successful attack demonstrations using imitation learning, strengthened by an ensemble Mixture-of-Experts architecture for robust generalization across scenarios. This expert policy then guides a DRL-based adversary through a KL-divergence regularization term. Due to the diversity of scenarios, expert policies may be imperfect. To address this, we further introduce a performance-aware annealing strategy that gradually reduces reliance on the expert as the adversary improves. Extensive experiments demonstrate that our method achieves outperforms existing approaches in terms of collision rate, attack efficiency, and training stability, especially in cases where the expert policy is sub-optimal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。