arXiv:2511.14592cs.ROcs.AI2025-11被引 2

首个综合评估驾驶内外安全风险的基准,揭示视觉语言模型在真实场景中的安全隐患。

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks

  • 构建统一评估框架,覆盖内外部共28个子类安全风险
  • 98K数据集微调后显著提升模型安全表现
  • 适合自动驾驶安全研究者与系统验证团队

视觉语言模型(VLMs)在自动驾驶中展现出巨大潜力,但其在安全关键场景下的适用性仍缺乏系统评估,引发安全担忧。这主要源于缺乏同时评估外部环境风险与车内驾驶行为安全的综合性基准。为填补这一空白,我们提出DSBench,首个面向自动驾驶安全的综合基准,可统一评估多种安全风险。该基准涵盖外部环境风险与车内驾驶行为安全两大类,细分为10个关键类别和28个子类别,全面覆盖多样化场景,确保对VLM在安全关键情境下表现的深入评估。对主流开源与闭源VLM的广泛测试显示,在复杂安全场景下性能显著下降,凸显迫切的安全隐患。为此,我们构建了包含98,000个实例的大规模数据集,聚焦内外部安全场景;实验证明,基于该数据集微调能显著提升现有VLM的安全能力,推动自动驾驶技术发展。基准工具包、代码及模型检查点将公开发布。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns. This issue arises from the lack of comprehensive benchmarks that assess both external environmental risks and in-cabin driving behavior safety simultaneously. To bridge this critical gap, we introduce DSBench, the first comprehensive Driving Safety Benchmark designed to assess a VLM's awareness of various safety risks in a unified manner. DSBench encompasses two major categories: external environmental risks and in-cabin driving behavior safety, divided into 10 key categories and a total of 28 sub-categories. This comprehensive evaluation covers a wide range of scenarios, ensuring a thorough assessment of VLMs' performance in safety-critical contexts. Extensive evaluations across various mainstream open-source and closed-source VLMs reveal significant performance degradation under complex safety-critical situations, highlighting urgent safety concerns. To address this, we constructed a large dataset of 98K instances focused on in-cabin and external safety scenarios, showing that fine-tuning on this dataset significantly enhances the safety performance of existing VLMs and paves the way for advancing autonomous driving technology. The benchmark toolkit, code, and model checkpoints will be publicly accessible.

自动驾驶视觉语言模型安全评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。