arXiv:2604.08212cs.CV2026-04被引 1

用专业指令微调让AI读懂道路病害,评估更准更省力。

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

  • 用27万+图像指令对训练专用模型,覆盖32类道路评估任务
  • 比现有模型在定位、推理和生成上提升超20%,符合行业标准
  • 适合交通部门用对话式工具替代多个专业系统

通用视觉语言模型在日常任务中表现良好,但在需要精确术语、结构化推理和遵循工程规范的专精领域表现不佳。本文探讨通过领域特定指令微调是否可实现全面的道路状况评估。研究构建了包含278,889个图像-指令-响应对的PaveInstruct数据集,整合了九个异构道路数据集的标注。基于该数据集训练的PaveGPT基础模型,在感知、理解与推理任务上对比当前最优视觉语言模型进行了评估。指令微调显著提升了模型能力,在空间定位、推理与生成任务上性能提升超过20%,且输出符合ASTM D6433标准。该成果使交通部门可部署统一的对话式评估工具,替代多个专用系统,简化流程并降低技术门槛。该方法为桥梁检测、铁路维护、建筑状况评估等基础设施领域提供了可扩展的指令驱动AI开发路径。

原文摘要 · Abstract (English)

General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring precise terminology, structured reasoning, and adherence to engineering standards. This work addresses whether domain-specific instruction tuning can enable comprehensive pavement condition assessment through vision-language models. PaveInstruct, a dataset containing 278,889 image-instruction-response pairs spanning 32 task types, was created by unifying annotations from nine heterogeneous pavement datasets. PaveGPT, a pavement foundation model trained on this dataset, was evaluated against state-of-the-art vision-language models across perception, understanding, and reasoning tasks. Instruction tuning transformed model capabilities, achieving improvements exceeding 20% in spatial grounding, reasoning, and generation tasks while producing ASTM D6433-compliant outputs. These results enable transportation agencies to deploy unified conversational assessment tools that replace multiple specialized systems, simplifying workflows and reducing technical expertise requirements. The approach establishes a pathway for developing instruction-driven AI systems across infrastructure domains including bridge inspection, railway maintenance, and building condition assessment.

道路评估视觉语言模型指令微调基础设施

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。