量化视觉语言模型可高效评估桥梁钢筋裸露损伤,兼顾准确与速度。
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
- 用不同精度量化LLaVA-1.5-7B模型分析254张钢筋裸露图像。
- Q5_K_M在质量、速度、资源消耗上综合表现最佳,优于其他量化级别。
- 适合部署在消费级显卡的桥梁智能巡检系统,尤其关注推理效率者适用。
桥梁结构检测是关键但耗人力的任务,需专家识别钢筋裸露、裂缝和腐蚀等损伤。本文系统研究了量化视觉语言模型(VLM)在自动化桥梁损伤评估中的应用,权衡描述质量、推理速度与资源开销。构建端到端流程:利用LLaVA-1.5-7B进行视觉损伤分析、结构化JSON提取及基于规则的优先级评分。为适配消费级GPU部署,对Q4_K_M、Q5_K_M、Q8_0三种量化级别在254张钢筋裸露图像上进行对比。引入五分制评估框架,涵盖损伤类型识别与严重程度分类。结果表明,Q5_K_M取得最优平衡:质量得分3.18±1.35/5.0,单图推理时间5.67秒,效率0.56质量/秒——比Q4_K_M高8.5%质量,仅慢4.5%;媲美Q8_0质量,但快25%。统计分析显示,Q5_K_M文本质量相关性最弱(-0.148),表明其性能不随描述长度波动,稳定可靠。
原文摘要 · Abstract (English)
Bridge infrastructure inspection is a critical but labor-intensive task requiring expert assessment of structural damage such as rebar exposure, cracking, and corrosion. This paper presents a comprehensive study of quantized Vision-Language Models (VLMs) for automated bridge damage assessment, focusing on the trade-offs between description quality, inference speed, and resource requirements. We develop an end-to-end pipeline combining LLaVA-1.5-7B for visual damage analysis, structured JSON extraction, and rule-based priority scoring. To enable deployment on consumer-grade GPUs, we conduct a systematic comparison of three quantization levels: Q4_K_M, Q5_K_M, and Q8\_0 across 254 rebar exposure images. We introduce a 5-point quality evaluation framework assessing damage type recognition, severity classification. Our results demonstrate that Q5_K_M achieves the optimal balance: quality score 3.18$\pm$1.35/5.0, inference time 5.67s/image, and 0.56 quality/sec efficiency -- 8.5% higher quality than Q4_K_M with only 4.5% speed reduction, while matching Q8_0's quality with 25% faster inference. Statistical analysis reveals Q5_K_M exhibits the weakest text-quality correlation (-0.148), indicating consistent performance regardless of description length.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。