arXiv:2411.08320cs.AI2024-11被引 8

评测大模型在工地安全领域的表现,发现其能帮上忙但需人工把关。

Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering

  • 用BCSP标准题库测试GPT-3.5和GPT-4o,评估其安全知识水平
  • GPT-4o准确率达84.6%,在安全管理与风险识别上表现较好
  • 提示词工程可提升性能13.5%,但无万能方案,需人工干预

建筑行业仍是高危领域。大语言模型(LLMs)的兴起为提升工作场所安全提供了新可能。然而,负责任地集成这些模型需系统评估,否则可能产生错误信息,导致误判,危及工人安全。本研究采用美国注册安全专业人员委员会(BCSP)的三套标准化考试,评估了GPT-3.5与GPT-4o在385道涵盖七个安全知识领域的题目上的表现。结果显示,两款模型均显著超过BCSP基准线:GPT-4o准确率达84.6%,GPT-3.5为73.8%。两者在安全管理体系与隐患识别控制方面表现良好,但在科学、数学、应急响应及防火防灾方面存在明显短板。错误分析揭示四大局限:知识缺失、推理缺陷、记忆不足与计算错误。提示词工程对性能影响显著,GPT-3.5最高波动达13.5%,GPT-4o为7.9%,但无单一配置通用有效。本研究从三方面推进认知:明确模型可辅助的安全领域与必须保留人工判断的环节;提供通过提示工程优化模型应用的实践洞见;为未来研发提供基于证据的方向。这些成果支持构建负责任的建筑安全智能管理,助力实现零伤亡目标。

原文摘要 · Abstract (English)

Construction remains one of the most hazardous sectors. Recent advancements in AI, particularly Large Language Models (LLMs), offer promising opportunities for enhancing workplace safety. However, responsible integration of LLMs requires systematic evaluation, as deploying them without understanding their capabilities and limitations risks generating inaccurate information, fostering misplaced confidence, and compromising worker safety. This study evaluates the performance of two widely used LLMs, GPT-3.5 and GPT-4o, across three standardized exams administered by the Board of Certified Safety Professionals (BCSP). Using 385 questions spanning seven safety knowledge areas, the study analyzes the models' accuracy, consistency, and reliability. Results show that both models consistently exceed the BCSP benchmark, with GPT-4o achieving an accuracy rate of 84.6% and GPT-3.5 reaching 73.8%. Both models demonstrate strengths in safety management systems and hazard identification and control, but exhibit weaknesses in science, mathematics, emergency response, and fire prevention. An error analysis identifies four primary limitations affecting LLM performance: lack of knowledge, reasoning flaws, memory issues, and calculation errors. Our study also highlights the impact of prompt engineering strategies, with variations in accuracy reaching 13.5% for GPT-3.5 and 7.9% for GPT-4o. However, no single prompt configuration proves universally effective. This research advances knowledge in three ways: by identifying areas where LLMs can support safety practices and where human oversight remains essential, by offering practical insights into improving LLM implementation through prompt engineering, and by providing evidence-based direction for future research and development. These contributions support the responsible integration of AI in construction safety management toward achieving zero injuries.

AI安全大模型施工安全提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。