arXiv:2410.00982cs.CV2024-10被引 22

提升视觉语言模型对交通危险事件的理解能力

ScVLM: Enhancing Vision-Language Model for Safety-Critical Event Understanding

  • 融合监督与对比学习,增强模型对事故视频的识别与描述能力
  • 在超过8600个事故数据上训练,显著减少模型幻觉现象
  • 适合自动驾驶和交通安全系统开发者使用

准确识别、理解并描述交通中的安全关键事件(SCEs),包括碰撞、轮胎撞击和近事故,对于高级驾驶辅助系统和自动驾驶系统至关重要。由于这些事件稀有,大多数通用视觉语言模型(VLMs)未能充分学习视频与叙事之间的关联,导致生成内容出现幻觉或遗漏关键安全特征。本文提出ScVLM,一种结合监督与对比学习的混合方法,用于分类SCE的严重程度与类型,并生成事件叙述。该方法通过分类任务提升VLM对驾驶视频的理解力,增强描述合理性。模型在包含超过8600个SCE的第二战略公路研究计划自然驾驶研究数据集(SHRP2 NDS)上进行训练与评估,结果表明其在生成上下文准确描述和抑制模型幻觉方面表现优异。代码将开源。

原文摘要 · Abstract (English)

Accurately identifying, understanding and describing traffic safety-critical events (SCEs), including crashes, tire strikes, and near-crashes, is crucial for advanced driver assistance systems, automated driving systems, and traffic safety. As SCEs are rare events, most general vision-language models (VLMs) have not been trained sufficiently to link SCE videos and narratives, which could lead to hallucinations and missing key safety characteristics. Here, we introduce ScVLM, a novel hybrid methodology that integrates supervised and contrastive learning techniques to classify the severity and types of SCEs, as well as to generate narrative descriptions of SCEs. This approach utilizes classification to enhance VLMs' comprehension of driving videos and improve the rationality of event descriptions. The proposed approach is trained on and evaluated by more than 8,600 SCEs from the Second Strategic Highway Research Program Naturalistic Driving Study dataset, the largest publicly accessible driving dataset with videos and SCE annotations. The results demonstrate the superiority of the proposed approach in generating contextually accurate event descriptions and mitigating VLM hallucinations. The code will be available at https://github.com/datadrivenwheels/ScVLM.

视觉语言模型交通安全自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。