arXiv:2511.22928cs.RO2025-11被引 4

用视觉语言模型提前识别驾驶中的潜在风险,提升自动驾驶安全预判能力。

Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models

  • 基于视觉语言模型,通过细微异常行为推理未显性暴露的风险。
  • 在自建数据集上,新模型比基线提升显著,能有效识别潜在风险链。
  • 适合研究自动驾驶安全预警与智能决策的学者和工程师。

确保安全性仍是自动驾驶车辆(AV)面临的关键挑战,尤其在罕见复杂场景中。一个关键但研究不足的方面是“潜在风险”——风险尚未显现,但可通过细微前兆(如异常行为或常识违背)推断。识别这些前兆需要强大的语义理解与推理能力,而现有自动驾驶或风险数据集中此类案例稀缺,且事故数据常缺少因果推理链标注,难以支持潜在风险识别。为此,我们提出PotentialRiskQA,一个面向潜在风险推理的新型多模态数据集,每条样本包含结构化场景描述、语义前兆与推断风险结果。基于该数据集,我们进一步设计了PR-Reasoner,一种专用于车载潜在风险推理的视觉语言模型框架。实验表明,在PotentialRiskQA上微调后,PR-Reasoner在潜在风险推理任务上显著优于基线视觉语言模型。整体而言,我们的数据集与模型为发展具备前瞻性与主动安全能力的自动驾驶系统提供了基础,推动更智能、更鲁棒的车辆演进。

原文摘要 · Abstract (English)

Ensuring safety remains a key challenge for autonomous vehicles (AVs), especially in rare and complex scenarios. One critical but understudied aspect is the \textbf{potential risk} situations, where the risk is \textbf{not yet observable} but can be inferred from subtle precursors, such as anomalous behaviors or commonsense violations. Recognizing these precursors requires strong semantic understanding and reasoning capabilities, which are often absent in current AV systems due to the scarcity of such cases in existing driving or risk-centric datasets. Moreover, current autonomous driving accident datasets often lack annotations of the causal reasoning chains behind incidents, which are essential for identifying potential risks before they become observable. To address these gaps, we introduce PotentialRiskQA, a novel vision-language dataset designed for reasoning about potential risks prior to observation. Each sample is annotated with structured scene descriptions, semantic precursors, and inferred risk outcomes. Based on this dataset, we further propose PR-Reasoner, a vision-language-model-based framework tailored for onboard potential risk reasoning. Experimental results show that fine-tuning on PotentialRiskQA enables PR-Reasoner to significantly enhance its performance on the potential risk reasoning task compared to baseline VLMs. Together, our dataset and model provide a foundation for developing autonomous systems with improved foresight and proactive safety capabilities, moving toward more intelligent and resilient AVs.

自动驾驶风险推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。