用视觉语言模型自动生成安全驾驶指令,提升行车风险识别能力。
Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach
- 融合前后摄像头视频,训练多模态模型生成驾驶建议
- 微调后模型能准确识别危险行为,正确率达78.3%
- 适合自动驾驶安全监控与智能车载系统研发者参考
大规模视觉语言模型(LVLMs)在需视觉理解的任务中表现优异,如目标检测,已在自动驾驶等领域展现应用潜力。例如,它们可对道路摄像头拍摄的视频生成安全导向描述。但要全面保障安全,还需监测驾驶员视角视频以识别使用手机等高危行为。因此,同步处理驾驶员与道路双视角视频输入成为必要。本研究构建了一个新数据集,并评估了LVLM在该任务上的表现。实验表明,预训练模型效果有限,而经过微调的模型可生成准确且具有安全意识的驾驶指令。然而,在识别细微或复杂事件方面仍存挑战。研究结果与错误分析为改进基于LVLM的安全驾驶系统提供了重要参考。
原文摘要 · Abstract (English)
Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous driving. For example, LVLMs can generate safety-oriented descriptions of videos captured by road-facing cameras. However, ensuring comprehensive safety requires monitoring driver-facing views as well to detect risky events, such as the use of mobiles while driving. Thus, the ability to process synchronized inputs is necessary from both driver-facing and road-facing cameras. In this study, we develop models and investigate the capabilities of LVLMs by constructing a dataset and evaluating their performance on this dataset. Our experimental results demonstrate that while pre-trained LVLMs have limited effectiveness, fine-tuned LVLMs can generate accurate and safety-aware driving instructions. Nonetheless, several challenges remain, particularly in detecting subtle or complex events in the video. Our findings and error analysis provide valuable insights that can contribute to the improvement of LVLM-based systems in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。