arXiv:2510.13108cs.CVcs.AI2025-10中稿 · ICRA被引 5

用视觉语言模型提升自动驾驶评估的上下文理解能力。

DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models

  • 基于视觉语言模型构建上下文感知的评估框架。
  • 在复杂场景中匹配人类偏好,准确率显著优于现有指标。
  • 适合关注自动驾驶安全与人性化评估的研究者。

当前自动驾驶规划器的评测仍难以对齐人类判断,现有指标如扩展预测驾驶员模型评分(EPDMS)在复杂场景中缺乏上下文感知能力。为此,本文提出DriveCritic框架,包含两个核心贡献:一是精心构建的DriveCritic数据集,涵盖需依赖上下文判断的挑战性场景,并附有成对的人类偏好标注;二是基于视觉语言模型(VLM)的DriveCritic评估模型。该模型通过两阶段监督与强化学习训练,能结合视觉与符号化上下文信息,对轨迹对进行评判。实验表明,DriveCritic在匹配人类偏好方面显著优于现有指标与基线方法,展现出强上下文感知能力。本工作为自动驾驶系统评估提供了更可靠、更符合人类直觉的基准。项目主页:https://song-jingyu.github.io/DriveCritic

原文摘要 · Abstract (English)

Benchmarking autonomous driving planners to align with human judgment remains a critical challenge, as state-of-the-art metrics like the Extended Predictive Driver Model Score (EPDMS) lack context awareness in nuanced scenarios. To address this, we introduce DriveCritic, a novel framework featuring two key contributions: the DriveCritic dataset, a curated collection of challenging scenarios where context is critical for correct judgment and annotated with pairwise human preferences, and the DriveCritic model, a Vision-Language Model (VLM) based evaluator. Fine-tuned using a two-stage supervised and reinforcement learning pipeline, the DriveCritic model learns to adjudicate between trajectory pairs by integrating visual and symbolic context. Experiments show DriveCritic significantly outperforms existing metrics and baselines in matching human preferences and demonstrates strong context awareness. Overall, our work provides a more reliable, human-aligned foundation to evaluating autonomous driving systems. The project page for DriveCritic is https://song-jingyu.github.io/DriveCritic

自动驾驶视觉语言模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。