用注意力机制生成高风险违规场景,提升自动驾驶测试全面性。
ROMAN: Reward-Orchestrated Multi-Head Attention Network for Autonomous Driving System Testing
- 通过多头注意力建模车辆与信号交互,生成复杂违规场景。
- 相比现有工具,违规场景数量提升55.96%,且覆盖全部交通法规条款。
- 适合自动驾驶安全验证团队,尤其关注法规合规性测试的场景。
自动驾驶系统(ADS)是自动驾驶车辆的大脑,关乎安全与效率。安全部署需在多样真实场景中充分测试,并遵守速度限制、信号服从及路权规则等交通法规。闯红灯、超速等违规行为带来严重安全隐患。然而,当前测试方法存在生成复杂高危违规场景能力有限,且难以处理多车交互与关键情境的问题。为此,我们提出ROMAN,一种结合多头注意力网络与交通法规权重机制的新型场景生成方法。多头注意力机制建模车辆、信号灯及其他因素间的交互关系;法规权重机制通过基于大模型的风险评估模块,从严重性和发生频率两个维度量化违规风险。我们在CARLA仿真平台中对百度Apollo ADS进行了测试,实验结果表明:与现有先进工具ABLE和LawBreaker相比,ROMAN平均违规数量分别高出7.91%和55.96%,同时保持更高场景多样性;更重要的是,仅有ROMAN成功生成了输入交通法规中每一条款对应的违规场景,从而发现更多高风险违规行为。
原文摘要 · Abstract (English)
Automated Driving System (ADS) acts as the brain of autonomous vehicles, responsible for their safety and efficiency. Safe deployment requires thorough testing in diverse real-world scenarios and compliance with traffic laws like speed limits, signal obedience, and right-of-way rules. Violations like running red lights or speeding pose severe safety risks. However, current testing approaches face significant challenges: limited ability to generate complex and high-risk law-breaking scenarios, and failing to account for complex interactions involving multiple vehicles and critical situations. To address these challenges, we propose ROMAN, a novel scenario generation approach for ADS testing that combines a multi-head attention network with a traffic law weighting mechanism. ROMAN is designed to generate high-risk violation scenarios to enable more thorough and targeted ADS evaluation. The multi-head attention mechanism models interactions among vehicles, traffic signals, and other factors. The traffic law weighting mechanism implements a workflow that leverages an LLM-based risk weighting module to evaluate violations based on the two dimensions of severity and occurrence. We have evaluated ROMAN by testing the Baidu Apollo ADS within the CARLA simulation platform and conducting extensive experiments to measure its performance. Experimental results demonstrate that ROMAN surpassed state-of-the-art tools ABLE and LawBreaker by achieving 7.91% higher average violation count than ABLE and 55.96% higher than LawBreaker, while also maintaining greater scenario diversity. In addition, only ROMAN successfully generated violation scenarios for every clause of the input traffic laws, enabling it to identify more high-risk violations than existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。