arXiv:2603.18178cs.CVcs.AI2026-03被引 2

用多模态模型提升自动驾驶中罕见事故的检测能力

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

  • 通过融合元数据、LLM描述和思维链监督,让视觉语言模型适应驾驶场景
  • 碰撞事件检测F1从0.00提升至0.69,准确率从35.35%升至77.27%
  • 适合关注自动驾驶安全与可解释推理的研究者和工程师

自车摄像头视频的快速增长给检测碰撞、近碰撞等安全关键事件带来挑战,这些场景短暂、稀有,通用视觉模型难以捕捉。尽管多模态大语言模型具备强泛化推理能力,但在驾驶场景中因领域和时间错位表现不佳。我们提出VLM-AutoDrive,一个模块化后训练框架,用于将预训练视觉语言模型(VLM)适配至高保真异常检测任务。该框架整合元数据生成的描述、LLM生成的文本、视觉问答对及思维链推理监督,实现领域对齐且可解释的学习。现成的VLM如NVIDIA Cosmos-Reason1 7B(CR1)在零样本设置下碰撞召回率接近0;经VLM-AutoDrive微调后,碰撞F1提升至0.69,整体准确率从35.35%增至77.27%。在真实世界Nexar摄像头视频上评估,该方法显著提升碰撞与近碰撞检测性能,并生成可解释的推理轨迹,弥合了自动驾驶中感知、因果与决策推理之间的差距。

原文摘要 · Abstract (English)

The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-collisions, scenarios that are brief, rare, and difficult for generic vision models to capture. While multimodal large language models (MLLMs) demonstrate strong general reasoning ability, they underperform in driving contexts due to domain and temporal misalignment. We introduce VLM-AutoDrive, a modular post-training framework for adapting pretrained Vision-Language Models (VLMs) to high-fidelity anomaly detection. The framework integrates metadata-derived captions, LLM-generated descriptions, visual question answering (VQA) pairs, and chain-of-thought (CoT) reasoning supervision to enable domain-aligned and interpretable learning. Off-the-shelf VLMs such as NVIDIA's Cosmos-Reason1 7B (CR1) exhibit near-zero Collision recall in zero-shot settings; fine-tuning with VLM-AutoDrive improves Collision F1 from 0.00 to 0.69 and overall accuracy from 35.35% to 77.27%. VLM-AutoDrive offers a scalable recipe for adapting general-purpose VLMs to safety-critical, temporally localized perception tasks. Evaluated on real-world Nexar dashcam videos, it achieves substantial gains in Collision and Near-Collision detection while producing interpretable reasoning traces, bridging the gap between perception, causality, and decision reasoning in autonomous driving.

自动驾驶视觉语言模型异常检测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。