用视觉语言模型自动识别城市道路缺陷并生成维修指令
InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
- 结合YOLO检测与VLM生成结构化维修方案
- 在公开数据集和实拍视频中准确识别多种缺陷
- 适合城市运维团队快速响应,提升巡检效率
智慧城市建设中,闭路电视(CCTV)网络日益用于基础设施监控。道路、桥梁和隧道出现裂缝、坑洼和液体泄漏等缺陷,威胁公共安全,需及时修复。人工巡检成本高且危险,现有自动化系统通常仅处理单一缺陷类型,或输出非结构化信息,无法直接指导维修人员。本文提出一个端到端框架,利用街景CCTV流,通过YOLO系列目标检测器实现多类缺陷检测与分割,并将结果输入视觉语言模型(VLM),生成包含事件描述、推荐工具、尺寸、修复方案和紧急提醒的结构化行动方案(JSON格式)。我们回顾了坑洼、裂缝和渗漏检测的研究进展,分析了QwenVL和LLaVA等大视觉语言模型的最新成果,介绍了早期原型设计。在公开数据集及实拍CCTV片段上的实验表明,该系统能准确识别多样缺陷并生成连贯摘要。最后讨论了向城市级部署扩展面临的挑战与方向。
原文摘要 · Abstract (English)
Infrastructure in smart cities is increasingly monitored by networks of closed circuit television (CCTV) cameras. Roads, bridges and tunnels develop cracks, potholes, and fluid leaks that threaten public safety and require timely repair. Manual inspection is costly and hazardous, and existing automatic systems typically address individual defect types or provide unstructured outputs that cannot directly guide maintenance crews. This paper proposes a comprehensive pipeline that leverages street CCTV streams for multi defect detection and segmentation using the YOLO family of object detectors and passes the detections to a vision language model (VLM) for scene aware summarization. The VLM generates a structured action plan in JSON format that includes incident descriptions, recommended tools, dimensions, repair plans, and urgent alerts. We review literature on pothole, crack and leak detection, highlight recent advances in large vision language models such as QwenVL and LLaVA, and describe the design of our early prototype. Experimental evaluation on public datasets and captured CCTV clips demonstrates that the system accurately identifies diverse defects and produces coherent summaries. We conclude by discussing challenges and directions for scaling the system to city wide deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。