用机器人自动巡检工地,自动生成安全报告。
Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- 分层架构:机器人导航+视觉语言模型+大模型协同工作
- 在模拟场景中召回率高,精度媲美顶尖闭源模型
- 全程可解释、人工可介入,适合工地安全自动化
建筑安全检查仍以人工为主,现有自动化方法依赖特定数据集,在快速变化的工地环境中难以维护。同时,机器人现场巡检仍需人工遥控和手动报告,成本高。本文提出一种多层框架,将机器人自主导航所见与通用工地安全规则关联,实现自动安全报告生成。系统包含两个核心模块:机器人侧通过SLAM与自主导航实现可重复覆盖和定点重访;AI侧由视觉语言模型(VLM)生成场景描述,检索组件结合OSHA及工地政策对描述进行锚定,另一VLM层基于规则评估安全状况,最后由大语言模型(LLM)整合输出安全报告。在模拟三种典型风险场景的实验室环境中验证,结果表明该方法具有高召回率,且精度优于或媲美现有闭源模型。本研究提供一个透明、可泛化的工作流,通过暴露各层中间结果并保持人工参与,突破黑箱模型局限,为后续扩展至更多任务和场景奠定基础。
原文摘要 · Abstract (English)
Construction safety inspection remains mostly manual, and automated approaches still rely on task-specific datasets that are hard to maintain in fast-changing construction environments due to frequent retraining. Meanwhile, field inspection with robots still depends on human teleoperation and manual reporting, which are labor-intensive. This paper aims to connect what a robot sees during autonomous navigation to the safety rules that are common in construction sites, automatically generating a safety inspection report. To this end, we proposed a multi-layer framework with two main modules: robotics and AI. On the robotics side, SLAM and autonomous navigation provide repeatable coverage and targeted revisits via waypoints. On AI side, a Vision Language Model (VLM)-based layer produces scene descriptions; a retrieval component powered grounds those descriptions in OSHA and site policies; Another VLM-based layer assesses the safety situation based on rules; and finally Large Language Model (LLM) layer generates safety reports based on previous outputs. The framework is validated with a proof-of-concept implementation and evaluated in a lab environment that simulates common hazards across three scenarios. Results show high recall with competitive precision compared to state-of-the-art closed-source models. This paper contributes a transparent, generalizable pipeline that moves beyond black-box models by exposing intermediate artifacts from each layer and keeping the human in the loop. This work provides a foundation for future extensions to additional tasks and settings within and beyond construction context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。