用大模型零样本评估路面状况,准确率超专家。
Zero-Shot Image-Based Large Language Model Approach to Road Pavement Monitoring
- 基于大模型图像识别与语言理解,零样本完成路面评估。
- 优化提示工程后,准确率超越人工专家,达90%以上。
- 适合城市级道路巡检,无需标注数据,部署成本低。
高效快速地评估路面状况对优先安排维护、保障交通安全及减少车辆磨损至关重要。传统人工检测存在主观性,现有机器学习方法依赖大量高质量标注数据,资源消耗大且难以适应不同路况。大型语言模型(LLM)的突破为解决这一问题提供了新可能。本文提出一种创新的自动化零样本学习方法,利用LLM的图像识别与自然语言理解能力,基于路面状况指数(PSCI)标准设计提示工程策略,构建多模型评估体系。通过与官方PSCI结果对比,选出最优模型。在谷歌街景图像上进行多层级专家对比测试,结果表明:经优化的模型采用结构化提示工程,在准确性与一致性上显著优于简单配置,甚至超越专家评估。该模型成功应用于街景图像,展示了其在城市级规模部署的潜力。研究凸显了LLM在自动化道路损伤评估中的变革性作用,强调了精细提示工程对可靠评估的关键影响。
原文摘要 · Abstract (English)
Effective and rapid evaluation of pavement surface condition is critical for prioritizing maintenance, ensuring transportation safety, and minimizing vehicle wear and tear. While conventional manual inspections suffer from subjectivity, existing machine learning-based methods are constrained by their reliance on large and high-quality labeled datasets, which require significant resources and limit adaptability across varied road conditions. The revolutionary advancements in Large Language Models (LLMs) present significant potential for overcoming these challenges. In this study, we propose an innovative automated zero-shot learning approach that leverages the image recognition and natural language understanding capabilities of LLMs to assess road conditions effectively. Multiple LLM-based assessment models were developed, employing prompt engineering strategies aligned with the Pavement Surface Condition Index (PSCI) standards. These models' accuracy and reliability were evaluated against official PSCI results, with an optimized model ultimately selected. Extensive tests benchmarked the optimized model against evaluations from various levels experts using Google Street View road images. The results reveal that the LLM-based approach can effectively assess road conditions, with the optimized model -employing comprehensive and structured prompt engineering strategies -outperforming simpler configurations by achieving high accuracy and consistency, even surpassing expert evaluations. Moreover, successfully applying the optimized model to Google Street View images demonstrates its potential for future city-scale deployments. These findings highlight the transformative potential of LLMs in automating road damage evaluations and underscore the pivotal role of detailed prompt engineering in achieving reliable assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。