融合车速定位与视觉数据,提升大模型对危险驾驶事件的识别能力。
Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

- 用低频视频+高频车载数据融合增强感知
- 仅用5000万参数就显著提升危险事件识别率
- 适合自动驾驶安全评估与智能交通研究者
多模态大模型在通用视觉理解上表现优异,但在高风险驾驶场景中仍难以准确感知和推理罕见的严重动态事件(如碰撞或近撞)。为此,我们提出一种新流程:将下采样视频帧与同步的高频率车况数据(IMU和GPS)及专用视觉模型的语义信息融合,生成高质量伪标签(含描述性标题和问答对),专门用于训练多模态大模型识别真实驾驶视频中的安全关键事件(SCEs)。实验表明,通过使用DoRA适配器微调开源模型QwenVL-2.5,仅需不到5000万可训练参数和有限算力,即可显著提升模型对安全关键事件的识别与解释能力。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, their application to safety-critical driving scenarios remains limited by an inability to accurately perceive and reason about rare high-stakes dynamic events, such as collisions or near-collisions. To address this, we introduce a pipeline that enhances MLLM perception by fusing downsampled video frames with synchronized high-frequency telematics data (IMU and GPS) and semantic insights from specialized computer vision models. Our pipeline generates high-quality pseudo-labels, including descriptive captions and question-answer pairs, specifically designed to train MLLMs to identify and describe Safety-Critical Events (SCEs) in real-world driving footage. We show the effectiveness of our approach fine-tuning the open-source QwenVL-2.5 model via DoRA adapters: our experiments demonstrate significant improvements in identifying and explaining safety-critical events, with fewer than 50M trainable parameters and limited computational budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。