构建1万张工地安全图数据集,验证大模型能否当安全检查员。
Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- 提出含三任务标注的ConstructionSite 10k数据集
- 大模型在零样本/少样本下表现良好但需微调
- 适合研究建筑安全检测的AI团队使用
建筑安全检查通常由人工现场识别安全隐患。随着视觉语言模型(VLMs)的发展,研究者开始探索其从现场图像中检测安全违规的能力。然而,缺乏公开数据集来全面评估和微调VLM在建筑安全检查中的表现。现有应用依赖小规模监督数据集,限制了其在未直接训练任务中的适用性。本文提出ConstructionSite 10k,包含10,000张工地图片,涵盖图像描述、安全规则违规视觉问答(VQA)和建筑元素视觉定位三个相互关联的任务。对当前最先进的大型预训练VLMs的评估显示,它们在零样本和少样本设置下具备显著泛化能力,但仍需额外训练才能适用于真实工地场景。该数据集为研究人员提供了训练和评估新架构与技术的宝贵基准。
原文摘要 · Abstract (English)
Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language Models (VLMs), researchers are exploring their use for tasks such as detecting safety rule violations from on-site images. However, there is a lack of open datasets to comprehensively evaluate and further fine-tune VLMs in construction safety inspection. Current applications of VLMs use small, supervised datasets, limiting their applicability in tasks they are not directly trained for. In this paper, we propose the ConstructionSite 10k, featuring 10,000 construction site images with annotations for three inter-connected tasks, including image captioning, safety rule violation visual question answering (VQA), and construction element visual grounding. Our subsequent evaluation of current state-of-the-art large pre-trained VLMs shows notable generalization abilities in zero-shot and few-shot settings, while additional training is needed to make them applicable to actual construction sites. This dataset allows researchers to train and evaluate their own VLMs with new architectures and techniques, providing a valuable benchmark for construction safety inspection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。