arXiv:2603.15364cs.AIcs.CL2026-03被引 1

用AI分析自动驾驶事故,找出64%的故障源于感知或规划问题。

CRASH: Cognitive Reasoning Agent for Safety Hazards in Autonomous Driving

  • 基于大模型构建智能代理,融合结构化与非结构化数据自动推理事故原因。
  • 分析2168起事故发现64%归因于感知或规划失败,近半为追尾碰撞。
  • 专家验证准确率达86%,适合安全研究者和自动驾驶开发者使用。

随着自动驾驶汽车系统日益复杂多样,识别运行故障的根本原因变得愈发困难。不同厂商在系统架构(从端到端到模块化)及算法、集成策略上的差异,限制了事故调查的标准化,阻碍了系统的安全性分析。本文分析了美国国家公路交通安全管理局(NHTSA)数据库中2021至2025年间的实际事故报告,构建了一个包含2,168起案例的数据集,覆盖超8000万英里行驶里程。为此提出CRASH(Cognitive Reasoning Agent for Safety Hazards),一个基于大语言模型的智能代理,可对事故报告中的标准字段与非结构化叙述进行联合推理。CRASH以统一表示形式生成摘要,判断主因并评估自动驾驶系统是否对事件有实质影响。结果表明:(1)CRASH将64%的事故归因于感知或规划失败,凸显基于推理分析在精准归因中的重要性;(2)约50%的事故为追尾碰撞,反映该问题在部署中仍长期存在。经五位领域专家验证,系统在归因准确性上达86%。整体上,CRASH展现出作为可扩展、可解释的自动化事故分析工具的巨大潜力,为安全研究与自动驾驶系统持续优化提供可行动洞察。

原文摘要 · Abstract (English)

As AVs grow in complexity and diversity, identifying the root causes of operational failures has become increasingly complex. The heterogeneity of system architectures across manufacturers, ranging from end-to-end to modular designs, together with variations in algorithms and integration strategies, limits the standardization of incident investigations and hinders systematic safety analysis. This work examines real-world AV incidents reported in the NHTSA database. We curate a dataset of 2,168 cases reported between 2021 and 2025, representing more than 80 million miles driven. To process this data, we introduce CRASH, Cognitive Reasoning Agent for Safety Hazards, an LLM-based agent that automates reasoning over crash reports by leveraging both standardized fields and unstructured narrative descriptions. CRASH operates on a unified representation of each incident to generate concise summaries, attribute a primary cause, and assess whether the AV materially contributed to the event. Our findings show that (1) CRASH attributes 64% of incidents to perception or planning failures, underscoring the importance of reasoning-based analysis for accurate fault attribution; and (2) approximately 50% of reported incidents involve rear-end collisions, highlighting a persistent and unresolved challenge in autonomous driving deployment. We further validate CRASH with five domain experts, achieving 86% accuracy in attributing AV system failures. Overall, CRASH demonstrates strong potential as a scalable and interpretable tool for automated crash analysis, providing actionable insights to support safety research and the continued development of autonomous driving systems.

自动驾驶事故分析大模型应用安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。