arXiv:2511.15720cs.AI2025-11

用图文模型自动识别工地隐患,提升施工安全监测效率。

Automated Hazard Detection in Construction Sites Using Large Language and Vision-Language Models

  • 融合文本与图像分析,利用大模型从报告和影像中提取隐患信息。
  • 在2.8万份事故报告上验证,小模型在特定提示下表现接近大模型。
  • 轻量级开源模型可低成本实现规则感知的安全监控,适合资源有限场景。

本论文探索一种多模态AI框架,通过联合分析文本与视觉数据,提升建筑工地安全水平。在事故数据常以报告、记录和现场图像等多格式存在的高风险环境中,传统方法难以有效整合信息。为此,论文提出结合大语言模型(LLMs)与视觉-语言模型(VLMs)的多模态框架,用于自动化识别工地安全隐患。两个案例研究评估了LLMs与VLMs的能力:第一项研究使用GPT 4o和GPT 4o mini,从28,000份OSHA事故报告(2000–2025年)中提取结构化洞察;第二项研究采用Molmo 7B和Qwen2 VL 2B两款轻量级开源视觉-语言模型,基于公开的ConstructionSite10k数据集,以自然语言提示检测规则级安全违规行为。该实验作为成本敏感的基准测试,支持大规模评估并具备真实标注。尽管模型规模较小,但在特定提示配置下,Molmo 7B与Qwen2 VL 2B展现出具有竞争力的性能,证实了低资源多模态系统在规则感知安全监控中的可行性。

原文摘要 · Abstract (English)

This thesis explores a multimodal AI framework for enhancing construction safety through the combined analysis of textual and visual data. In safety-critical environments such as construction sites, accident data often exists in multiple formats, such as written reports, inspection records, and site imagery, making it challenging to synthesize hazards using traditional approaches. To address this, this thesis proposed a multimodal AI framework that combines text and image analysis to assist in identifying safety hazards on construction sites. Two case studies were consucted to evaluate the capabilities of large language models (LLMs) and vision-language models (VLMs) for automated hazard identification.The first case study introduces a hybrid pipeline that utilizes GPT 4o and GPT 4o mini to extract structured insights from a dataset of 28,000 OSHA accident reports (2000-2025). The second case study extends this investigation using Molmo 7B and Qwen2 VL 2B, lightweight, open-source VLMs. Using the public ConstructionSite10k dataset, the performance of the two models was evaluated on rule-level safety violation detection using natural language prompts. This experiment served as a cost-aware benchmark against proprietary models and allowed testing at scale with ground-truth labels. Despite their smaller size, Molmo 7B and Quen2 VL 2B showed competitive performance in certain prompt configurations, reinforcing the feasibility of low-resource multimodal systems for rule-aware safety monitoring.

工地安全多模态视觉语言模型自动化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。