用专业指令微调让AI读懂道路病害,评估更准更省力。
Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment
- 用27万+图像指令对训练专用模型,覆盖32类道路评估任务
- 比现有模型在定位、推理和生成上提升超20%,符合行业标准
- 适合交通部门用对话式工具替代多个专业系统
通用视觉语言模型在日常任务中表现良好,但在需要精确术语、结构化推理和遵循工程规范的专精领域表现不佳。本文探讨通过领域特定指令微调是否可实现全面的道路状况评估。研究构建了包含278,889个图像-指令-响应对的PaveInstruct数据集,整合了九个异构道路数据集的标注。基于该数据集训练的PaveGPT基础模型,在感知、理解与推理任务上对比当前最优视觉语言模型进行了评估。指令微调显著提升了模型能力,在空间定位、推理与生成任务上性能提升超过20%,且输出符合ASTM D6433标准。该成果使交通部门可部署统一的对话式评估工具,替代多个专用系统,简化流程并降低技术门槛。该方法为桥梁检测、铁路维护、建筑状况评估等基础设施领域提供了可扩展的指令驱动AI开发路径。
原文摘要 · Abstract (English)
General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring precise terminology, structured reasoning, and adherence to engineering standards. This work addresses whether domain-specific instruction tuning can enable comprehensive pavement condition assessment through vision-language models. PaveInstruct, a dataset containing 278,889 image-instruction-response pairs spanning 32 task types, was created by unifying annotations from nine heterogeneous pavement datasets. PaveGPT, a pavement foundation model trained on this dataset, was evaluated against state-of-the-art vision-language models across perception, understanding, and reasoning tasks. Instruction tuning transformed model capabilities, achieving improvements exceeding 20% in spatial grounding, reasoning, and generation tasks while producing ASTM D6433-compliant outputs. These results enable transportation agencies to deploy unified conversational assessment tools that replace multiple specialized systems, simplifying workflows and reducing technical expertise requirements. The approach establishes a pathway for developing instruction-driven AI systems across infrastructure domains including bridge inspection, railway maintenance, and building condition assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。