arXiv:2502.20572cs.CVcs.CL2025-02被引 2

小模型实现实时交通危险检测,适合边缘设备部署。

HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices

  • 微调Qwen2-VL-2B构建轻量视觉语言模型,支持边缘推理。
  • 在HazardQA数据集上F1分数提升89%,接近GPT-4o表现。
  • 专为交通危险场景设计,适合智能交通系统开发者使用。

交通安全隐患在现代城市中日益突出,主要源于车辆增多与道路网络复杂化。传统安全事件检测系统依赖传感器和传统机器学习算法,需大量数据采集与复杂训练流程。本文提出HazardNet,一个基于预训练Qwen2-VL-2B(20亿参数)微调的小规模视觉语言模型,利用先进多模态推理能力提升交通安全性。同时构建了专用于安全关键事件的HazardQA视觉问答数据集。实验表明,微调后的HazardNet在F1分数上相较基线模型最高提升89%,在部分任务中性能接近甚至超过更大的GPT-4o模型。该成果展示了在边缘设备上实现高效、实时交通危险检测的潜力,有助于降低事故率并优化城市交通管理。HazardNet模型与HazardQA数据集已公开于Hugging Face平台。

原文摘要 · Abstract (English)

Traffic safety remains a vital concern in contemporary urban settings, intensified by the increase of vehicles and the complicated nature of road networks. Traditional safety-critical event detection systems predominantly rely on sensor-based approaches and conventional machine learning algorithms, necessitating extensive data collection and complex training processes to adhere to traffic safety regulations. This paper introduces HazardNet, a small-scale Vision Language Model designed to enhance traffic safety by leveraging the reasoning capabilities of advanced language and vision models. We built HazardNet by fine-tuning the pre-trained Qwen2-VL-2B model, chosen for its superior performance among open-source alternatives and its compact size of two billion parameters. This helps to facilitate deployment on edge devices with efficient inference throughput. In addition, we present HazardQA, a novel Vision Question Answering (VQA) dataset constructed specifically for training HazardNet on real-world scenarios involving safety-critical events. Our experimental results show that the fine-tuned HazardNet outperformed the base model up to an 89% improvement in F1-Score and has comparable results with improvement in some cases reach up to 6% when compared to larger models, such as GPT-4o. These advancements underscore the potential of HazardNet in providing real-time, reliable traffic safety event detection, thereby contributing to reduced accidents and improved traffic management in urban environments. Both HazardNet model and the HazardQA dataset are available at https://huggingface.co/Tami3/HazardNet and https://huggingface.co/datasets/Tami3/HazardQA, respectively.

边缘计算视觉语言模型交通安全小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。