用视觉语言模型提升自动驾驶评估的可解释性与场景感知能力
DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models

- 结合规则引擎与视觉语言模型,先理解环境再判断驾驶行为
- 在3.3万条复杂驾驶样本上测试,评估准确率提升21.23%和6.5%
- 适合需要可解释性评估的自动驾驶研发与安全验证团队
自动驾驶正转向端到端策略学习,但驾驶质量评估面临可解释性与上下文感知的挑战。传统基于规则的指标(如EPDMS)虽易理解却缺乏场景感知,而现有基于视觉语言模型(VLM)的评估则因输出模糊且物理基础薄弱而受限。为此,我们提出DriveJudge:一个融合规则驱动评估与VLM推理的驾驶评价代理。它先通过VLM解析环境上下文,再选择性调用物理可信的确定性规则函数进行判断。为训练与评估,我们构建了包含33,577个高难度驾驶样本的大规模数据集,并引入两个符合人类判断的基准任务:驾驶质量分类与轨迹偏好选择。DriveJudge在驾驶质量分类上比EPDMS高出21.23 AUC,优于近期VLM方法DriveCritic在轨迹偏好选择中6.5%,树立了可解释、精准驾驶评估的新标准。
原文摘要 · Abstract (English)
Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-dependent. Commonly used rule-based driving metrics like EPDMS are interpretable but lack context-awareness, while recent VLMbased evaluations are context-aware but limited by ambiguous VLM outputs and weak physical grounding. To evaluate driving in a manner that is both interpretable and context-aware, we introduce DriveJudge. DriveJudge is a driving evaluation agent that combines rule-grounded evaluation with Vision-Language Model (VLM) reasoning and selectively invokes physically-grounded deterministic rule functions after interpreting the environmental context. To train and evaluate DriveJudge, we curate a large-scale dataset of 33,577 challenging driving samples with human annotations on whether the driving behavior is reasonable in the given scenario. With this dataset, we address the underexplored problem of driving metric evaluation, and introduce two human-aligned benchmark tasks: Driving Quality Classification and Trajectory Preference Selection. DriveJudge outperforms EPDMS for driving quality classification by 21.23 AUC, and the recent VLM-based DriveCritic for trajectory preference selection by 6.5%, setting a new standard for interpretable and precise driving evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。