arXiv:2601.12138cs.AI2026-01

为车载大模型设计安全风险分类体系,识别129种驾驶场景中的潜在危险

DriveSafe: A Hierarchical Risk Taxonomy for Safety-Critical LLM-Based Driving Assistants

  • 构建四级分层风险分类体系,覆盖技术、法律、社会与伦理维度
  • 发现六款主流大模型在危险驾驶请求前拒绝率不足,通用安全对齐失效
  • 基于真实交通法规和专家评审,适合自动驾驶与智能座舱安全研究者

大型语言模型(LLMs)正越来越多地被集成到车载数字助手中,不当、模糊或违法的响应可能引发严重安全、伦理和监管后果。尽管对大模型安全的关注日益增加,现有分类体系和评估框架仍以通用为主,难以捕捉真实驾驶场景中的特定风险。本文提出DriveSafe,一个四层结构的层级化风险分类体系,系统刻画基于大模型的驾驶助手的安全关键失效模式。该分类体系包含129个细粒度原子风险类别,涵盖技术、法律、社会及伦理维度,基于真实交通法规与安全原则,并经领域专家评审。为验证所构建提示的安全相关性与现实性,我们在六款广泛应用的大模型上评估其拒绝行为。分析表明,这些模型在面对不安全或不合规的驾驶相关查询时,往往未能适当拒绝,凸显通用安全对齐在驾驶场景中的局限性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly integrated into vehicle-based digital assistants, where unsafe, ambiguous, or legally incorrect responses can lead to serious safety, ethical, and regulatory consequences. Despite growing interest in LLM safety, existing taxonomies and evaluation frameworks remain largely general-purpose and fail to capture the domain-specific risks inherent to real-world driving scenarios. In this paper, we introduce DriveSafe, a hierarchical, four-level risk taxonomy designed to systematically characterize safety-critical failure modes of LLM-based driving assistants. The taxonomy comprises 129 fine-grained atomic risk categories spanning technical, legal, societal, and ethical dimensions, grounded in real-world driving regulations and safety principles and reviewed by domain experts. To validate the safety relevance and realism of the constructed prompts, we evaluate their refusal behavior across six widely deployed LLMs. Our analysis shows that the evaluated models often fail to appropriately refuse unsafe or non-compliant driving-related queries, underscoring the limitations of general-purpose safety alignment in driving contexts.

大模型安全驾驶助手风险分类LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。