arXiv:2410.12225cs.CV2024-10被引 3

用视觉语言模型零样本检测工地安全帽,提升施工安全监控效率

Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety

  • 构建级联检测流程,利用OWLv2模型实现零样本识别安全帽
  • 在5210张真实工地图像上达到0.6493的平均精度
  • 提出新数据集并揭示当前模型在安全场景中的局限性

本文评估了视觉语言模型(VLMs)在零样本检测与关联工地安全帽方面的应用,以提升施工安全。鉴于建筑工地头部受伤风险高,正确佩戴安全帽至关重要。研究探讨了基础模型OWLv2在真实工地图像中检测安全帽的适用性。主要贡献包括:通过筛选和整合现有数据集,构建了新的基准数据集Hardhat Safety Detection Dataset;开发了一种级联检测方法。在5,210张图像上的实验表明,OWLv2模型在安全帽检测任务中平均精度达0.6493。进一步分析了实际应用中的局限性及改进方向,揭示了当前基础模型在安全感知领域中的优劣。

原文摘要 · Abstract (English)

This paper evaluates the use of vision-language models (VLMs) for zero-shot detection and association of hardhats to enhance construction safety. Given the significant risk of head injuries in construction, proper enforcement of hardhat use is critical. We investigate the applicability of foundation models, specifically OWLv2, for detecting hardhats in real-world construction site images. Our contributions include the creation of a new benchmark dataset, Hardhat Safety Detection Dataset, by filtering and combining existing datasets and the development of a cascaded detection approach. Experimental results on 5,210 images demonstrate that the OWLv2 model achieves an average precision of 0.6493 for hardhat detection. We further analyze the limitations and potential improvements for real-world applications, highlighting the strengths and weaknesses of current foundation models in safety perception domains.

视觉语言模型零样本检测工地安全目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。