构建首个驾驶安全认知评估基准,提升视觉语言模型在自动驾驶中的安全判断能力。
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
- 提出SCD-Bench评估框架,专用于测试VLM在交互式驾驶场景下的安全认知能力。
- 构建32.4万条高质量数据集SCD-Training,使模型在多个基准上显著提升。
- 采用半自动标注与大模型评估,实现高效可靠的安全性评测,适合自动驾驶研究者。
确保视觉语言模型(VLMs)在自动驾驶系统中的安全性至关重要,但现有研究多集中于常规基准而非安全关键评估。本文提出SCD-Bench(安全认知驾驶基准),一个专为评估VLM在交互式驾驶场景中安全认知能力而设计的新框架。为解决数据标注的可扩展性问题,引入ADA(自动驾驶标注)系统,结合领域专家审核进行优化。同时提出基于大语言模型的自动化评估流程,与人工专家判断一致率达98%以上。为应对模型与驾驶安全认知对齐的挑战,构建了首个大规模数据集SCD-Training,包含324.35万条高质量样本。大量实验表明,基于SCD-Training训练的模型不仅在SCD-Bench上表现优异,也在通用和领域特定基准上获得显著提升,为增强视觉语言系统在自动驾驶中的安全交互提供了新思路。
原文摘要 · Abstract (English)
Ensuring the safety of vision-language models (VLMs) in autonomous driving systems is of paramount importance, yet existing research has largely focused on conventional benchmarks rather than safety-critical evaluation. In this work, we present SCD-Bench (Safety Cognition Driving Benchmark) a novel framework specifically designed to assess the safety cognition capabilities of VLMs within interactive driving scenarios. To address the scalability challenge of data annotation, we introduce ADA (Autonomous Driving Annotation), a semi-automated labeling system, further refined through expert review by professionals with domain-specific knowledge in autonomous driving. To facilitate scalable and consistent evaluation, we also propose an automated assessment pipeline leveraging large language models, which demonstrates over 98% agreement with human expert judgments. In addressing the broader challenge of aligning VLMs with safety cognition in driving environments, we construct SCD-Training, the first large-scale dataset tailored for this task, comprising 324.35K high-quality samples. Through extensive experiments, we show that models trained on SCD-Training exhibit marked improvements not only on SCD-Bench, but also on general and domain-specific benchmarks, offering a new perspective on enhancing safety-aware interactions in vision-language systems for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。