arXiv:2602.16073cs.ROcs.AI2026-02中稿 · 2026 IEEE Intellig…被引 1

构建可评估多目标优先级驾驶行为的基准,提升自动驾驶系统测试真实性。

ScenicRules: An Autonomous Driving Benchmark with Multi-Objective Specifications and Abstract Scenarios

  • 用分层规则框架形式化多个驾驶目标及优先级关系
  • 在随机环境中验证,模型失败率显著暴露于优先级冲突场景
  • 适合自动驾驶安全评估与决策算法研究者使用

开发复杂交通环境下的自动驾驶系统需兼顾多重目标,如避障、遵守交通规则和高效前行。在许多情况下,这些目标无法同时满足,自然产生优先级关系。驾驶规则依赖上下文,因此需形式化建模规则适用的环境场景。现有自动驾驶评测基准缺乏多目标优先规则与形式化环境模型的结合。本文提出ScenicRules,一个在随机环境中基于优先级多目标规范评估自动驾驶系统的基准。首先形式化一组多样化目标作为量化评估指标;其次设计分层规则书框架,以可解释且可调节的方式编码多个目标及其优先级;最后构建一套紧凑而具代表性的场景集合,涵盖多样驾驶情境与近事故情况,均以Scenic语言形式化建模。实验表明,所形式化的目标与分层规则书与人类驾驶判断高度一致,且本基准能有效暴露智能体在优先级目标上的失败。该基准可在https://github.com/BerkeleyLearnVerify/ScenicRules/ 获取。

原文摘要 · Abstract (English)

Developing autonomous driving systems for complex traffic environments requires balancing multiple objectives, such as avoiding collisions, obeying traffic rules, and making efficient progress. In many situations, these objectives cannot be satisfied simultaneously, and explicit priority relations naturally arise. Also, driving rules require context, so it is important to formally model the environment scenarios within which such rules apply. Existing benchmarks for evaluating autonomous vehicles lack such combinations of multi-objective prioritized rules and formal environment models. In this work, we introduce ScenicRules, a benchmark for evaluating autonomous driving systems in stochastic environments under prioritized multi-objective specifications. We first formalize a diverse set of objectives to serve as quantitative evaluation metrics. Next, we design a Hierarchical Rulebook framework that encodes multiple objectives and their priority relations in an interpretable and adaptable manner. We then construct a compact yet representative collection of scenarios spanning diverse driving contexts and near-accident situations, formally modeled in the Scenic language. Experimental results show that our formalized objectives and Hierarchical Rulebooks align well with human driving judgments and that our benchmark effectively exposes agent failures with respect to the prioritized objectives. Our benchmark can be accessed at https://github.com/BerkeleyLearnVerify/ScenicRules/.

自动驾驶多目标评估规则建模基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。