构建复杂交通规则推理基准,测试自动驾驶模型真实场景下的规则理解能力。
DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving
- 设计五层认知阶梯,从单规则到多规则冲突解决分阶段评估。
- 14个主流模型在规则冲突场景下性能显著下降,暴露短板。
- 适合研究自动驾驶感知与决策的学者,尤其关注规则推理的团队。
多模态大语言模型正快速成为端到端自动驾驶系统的核心智能引擎。当前关键挑战在于评估这些模型是否真正理解并遵守复杂的现实交通规则。然而,现有基准多聚焦于单一规则任务如交通标志识别,忽略了真实驾驶中多规则并发与冲突的复杂性。导致模型在简单任务表现良好,但在复杂场景常出错甚至违规。为此,我们提出DriveCombo——一个基于文本与视觉的组合式交通规则推理基准。受人类驾驶员认知发展启发,我们设计了系统的五层认知阶梯,评估从单规则理解到多规则整合与冲突解决的能力,实现跨认知阶段的量化评估。我们进一步提出Rule2Scene Agent,通过规则构造与场景生成,将语言描述的交通规则映射为动态驾驶场景,支持场景级交通规则视觉推理。对14个主流多模态大模型的评估显示,随着任务复杂度增加,模型性能显著下降,尤其在规则冲突时。在拆分数据集并进行训练集微调后,模型在交通规则推理及下游规划能力上均有显著提升。结果表明DriveCombo能有效推动合规且智能的自动驾驶系统发展。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are rapidly becoming the intelligence brain of end-to-end autonomous driving systems. A key challenge is to assess whether MLLMs can truly understand and follow complex real-world traffic rules. However, existing benchmarks mainly focus on single-rule scenarios like traffic sign recognition, neglecting the complexity of multi-rule concurrency and conflicts in real driving. Consequently, models perform well on simple tasks but often fail or violate rules in real world complex situations. To bridge this gap, we propose DriveCombo, a text and vision-based benchmark for compositional traffic rule reasoning. Inspired by human drivers' cognitive development, we propose a systematic Five-Level Cognitive Ladder that evaluates reasoning from single-rule understanding to multi-rule integration and conflict resolution, enabling quantitative assessment across cognitive stages. We further propose a Rule2Scene Agent that maps language-based traffic rules to dynamic driving scenes through rule crafting and scene generation, enabling scene-level traffic rule visual reasoning. Evaluations of 14 mainstream MLLMs reveal performance drops as task complexity grows, particularly during rule conflicts. After splitting the dataset and fine-tuning on the training set, we further observe substantial improvements in both traffic rule reasoning and downstream planning capabilities. These results highlight the effectiveness of DriveCombo in advancing compliant and intelligent autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。