arXiv:2506.07436cs.CVcs.AI2025-06被引 2

对比五款大模型在工地安全识别中的表现,发现提示词设计影响关键结果。

Prompt to Protection: A Comparative Study of Multimodal LLMs in Construction Hazard Recognition

  • 用三种提示策略测试大模型对工地图像的隐患识别能力
  • 思维链提示使各模型准确率普遍提升,GPT-4.5和GPT-o3表现最优
  • 提示工程是提升模型可靠性的重要手段,适合安全应用落地

多模态大语言模型(LLM)为建筑工地视觉隐患识别带来新可能。与依赖特定领域训练的传统计算机视觉模型不同,现代LLM可通过自然语言提示理解复杂场景。然而,其在施工安全等关键任务中的表现尚缺乏系统评估。本研究对比了五种前沿模型:Claude-3 Opus、GPT-4.5、GPT-4o、GPT-o3 和 Gemini 2.0 Pro,采用零样本、少样本和思维链(CoT)三种提示策略,在真实工地图像上评估隐患识别能力。使用精确率、召回率和F1分数进行定量分析。结果表明,提示策略显著影响性能,思维链提示在所有模型中均提升准确率;其中GPT-4.5与GPT-o3在多数条件下表现最佳。研究揭示提示设计在提升多模态大模型安全性应用中的核心作用,为实际部署提供可操作建议。

原文摘要 · Abstract (English)

The recent emergence of multimodal large language models (LLMs) has introduced new opportunities for improving visual hazard recognition on construction sites. Unlike traditional computer vision models that rely on domain-specific training and extensive datasets, modern LLMs can interpret and describe complex visual scenes using simple natural language prompts. However, despite growing interest in their applications, there has been limited investigation into how different LLMs perform in safety-critical visual tasks within the construction domain. To address this gap, this study conducts a comparative evaluation of five state-of-the-art LLMs: Claude-3 Opus, GPT-4.5, GPT-4o, GPT-o3, and Gemini 2.0 Pro, to assess their ability to identify potential hazards from real-world construction images. Each model was tested under three prompting strategies: zero-shot, few-shot, and chain-of-thought (CoT). Zero-shot prompting involved minimal instruction, few-shot incorporated basic safety context and a hazard source mnemonic, and CoT provided step-by-step reasoning examples to scaffold model thinking. Quantitative analysis was performed using precision, recall, and F1-score metrics across all conditions. Results reveal that prompting strategy significantly influenced performance, with CoT prompting consistently producing higher accuracy across models. Additionally, LLM performance varied under different conditions, with GPT-4.5 and GPT-o3 outperforming others in most settings. The findings also demonstrate the critical role of prompt design in enhancing the accuracy and consistency of multimodal LLMs for construction safety applications. This study offers actionable insights into the integration of prompt engineering and LLMs for practical hazard recognition, contributing to the development of more reliable AI-assisted safety systems.

多模态安全识别提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。