arXiv:2508.02073cs.AI2025-08被引 1

用外部案例库提升大模型施工风险识别能力,无需微调

Large model retrieval enhancement framework for construction site risk identification

  • 通过提示词调优融合检索到的相似案例和外部知识
  • 使GLM-4V准确率提升至50%,比基线高35.49%
  • 适合需要快速部署、通用性强的智能安全检测场景

本研究针对施工场地危险识别问题,提出一种无需微调的检索增强框架,以提升大语言模型(LLMs)性能。现有基于LLM的方法存在图像文本匹配困难、指令调优泛化性差且资源消耗大等问题。本文方法通过提示词调优动态整合外部知识与检索到的相似案例,克服了大模型在领域知识和特征关联上的局限。框架包含案例数据库、图像检索模块与基于LLM的推理模块。在真实工地数据上评估显示,该方法将GLM-4V的准确率提升至50%,相较基线提升35.49%,各类风险类型均表现稳定。消融实验验证了图像检索策略的有效性,表明基于LPIPS与CLIP的方法更具优势。该技术显著提升识别准确率与上下文理解能力,具备强泛化性,为建筑工地智能安全风险检测提供了可行路径。

原文摘要 · Abstract (English)

This study addresses construction site hazard identification by proposing a retrieval-augmented framework that enhances large language models (LLMs) without requiring fine-tuning. Current LLM-based approaches face limitations: image-text matching struggles with complex hazards, while instruction tuning lacks generalization and is resource-intensive. Our method dynamically integrates external knowledge and retrieved similar cases via prompt tuning, overcoming LLMs' limitations in domain knowledge and feature correlation. The framework comprises a case database, an image retrieval module, and an LLM-based reasoning module. Evaluated on real-site data, our approach boosted GLM-4V's accuracy to 50%, a 35.49% improvement over baselines, with consistent gains across hazard types. Ablation studies validated the effectiveness of our image retrieval strategy, showing the superiority of our LPIPS- and CLIP-based method. The proposed technique significantly improves identification accuracy and contextual understanding, demonstrating strong generalization and offering a practical path for intelligent safety risk detection in construction.

风险识别大模型检索增强施工安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。