用多模态AI自动生成工地安全检查报告,提升效率与准确性。
Automating construction safety inspections using a multi-modal vision-language RAG framework
- 结合视觉与音频信息,通过检索增强生成框架实现智能分析。
- 在真实数据上达到F1分数0.82,召回率0.96,优于单模态模型。
- 适合建筑安全监管、智能工地系统开发者使用。
传统施工安全检查方法效率低下,需处理大量信息。近年来大型视觉语言模型(LVLM)的发展为自动化安全检查提供了可能,但现有应用存在响应无关、输入模式受限及幻觉等问题。利用大语言模型(LLM)时,训练数据不足且缺乏实时适应性。本研究提出SiteShield,一种基于多模态LVLM的检索增强生成(RAG)框架,通过整合视觉与音频输入,实现工地安全检查报告的自动化生成。基于真实数据测试显示,SiteShield在无RAG的单模态LLM基础上,取得F1分数0.82、汉明损失0.04、精确率0.76、召回率0.96的性能表现。结果表明,SiteShield为提升安全报告的信息检索与生成效率提供了新路径。
原文摘要 · Abstract (English)
Conventional construction safety inspection methods are often inefficient as they require navigating through large volume of information. Recent advances in large vision-language models (LVLMs) provide opportunities to automate safety inspections through enhanced visual and linguistic understanding. However, existing applications face limitations including irrelevant or unspecific responses, restricted modal inputs and hallucinations. Utilisation of Large Language Models (LLMs) for this purpose is constrained by availability of training data and frequently lack real-time adaptability. This study introduces SiteShield, a multi-modal LVLM-based Retrieval-Augmented Generation (RAG) framework for automating construction safety inspection reports by integrating visual and audio inputs. Using real-world data, SiteShield outperformed unimodal LLMs without RAG with an F1 score of 0.82, hamming loss of 0.04, precision of 0.76, and recall of 0.96. The findings indicate that SiteShield offers a novel pathway to enhance information retrieval and efficiency in generating safety reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。