用多智能体强化学习生成逼真危险场景,提升自动驾驶系统安全测试效率。
MARL-OT: Multi-Agent Reinforcement Learning Guided Online Fuzzing to Detect Safety Violation in Autonomous Driving Systems
- 基于多智能体强化学习动态生成协同威胁车辆行为
- 相比最先进方法,安全漏洞检测率提升136.2%
- 适合自动驾驶安全验证与仿真测试研究人员
自动驾驶系统(ADS)是安全关键型系统,真实世界中的安全违规可能造成重大损失。部署前需进行严格测试,其中仿真测试起关键作用。然而,ADS通常结构复杂,包含感知、规划等多模块或端到端训练模型。传统离线方法如遗传算法(GA)仅能生成预设轨迹,因进化特性难以高效触发不同场景下的安全违规。在线方法如单智能体强化学习(RL)可实时调整动态轨迹以适应场景,但难以捕捉多车交互产生的复杂边缘情况。多智能体强化学习(MARL)在协作任务中表现优异,但面临收敛难题。本文提出MARL-OT,一个可扩展的框架,利用MARL对周围车辆行为进行高层引导,触发多种危险场景,供基于规则的在线模糊测试器探索潜在的安全违规,从而生成动态且真实的危险场景。实验表明,该方法相较当前最先进(SOTA)测试技术,安全违规检测率提升最高达136.2%。
原文摘要 · Abstract (English)
Autonomous Driving Systems (ADSs) are safety-critical, as real-world safety violations can result in significant losses. Rigorous testing is essential before deployment, with simulation testing playing a key role. However, ADSs are typically complex, consisting of multiple modules such as perception and planning, or well-trained end-to-end autonomous driving systems. Offline methods, such as the Genetic Algorithm (GA), can only generate predefined trajectories for dynamics, which struggle to cause safety violations for ADSs rapidly and efficiently in different scenarios due to their evolutionary nature. Online methods, such as single-agent reinforcement learning (RL), can quickly adjust the dynamics' trajectory online to adapt to different scenarios, but they struggle to capture complex corner cases of ADS arising from the intricate interplay among multiple vehicles. Multi-agent reinforcement learning (MARL) has a strong ability in cooperative tasks. On the other hand, it faces its own challenges, particularly with convergence. This paper introduces MARL-OT, a scalable framework that leverages MARL to detect safety violations of ADS resulting from surrounding vehicles' cooperation. MARL-OT employs MARL for high-level guidance, triggering various dangerous scenarios for the rule-based online fuzzer to explore potential safety violations of ADS, thereby generating dynamic, realistic safety violation scenarios. Our approach improves the detected safety violation rate by up to 136.2% compared to the state-of-the-art (SOTA) testing technique.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。