arXiv:2604.02804cs.CVcs.AI2026-04

构建首个面向道路病害的多轮交互式视觉语言评测基准

PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis

  • 设计支持多任务统一评测的公路病害感知框架
  • 包含真实图像与多轮问答数据,覆盖识别、定位与养护推理
  • 适合智能交通、自动驾驶及城市运维领域研究者使用

路面状况评估对道路安全与维护至关重要。现有研究多集中于分类、检测、分割等传统计算机视觉任务,但实际巡检还需定量分析、解释与交互决策支持。当前数据集普遍局限于单模态感知,缺乏多轮交互、事实依据推理及感知与视觉语言分析的关联。为此,我们提出PaveBench,一个大规模公路病害感知与交互式视觉语言分析基准。该基准支持四大核心任务:分类、目标检测、语义分割与视觉语言问答。提供统一任务定义与评估协议。视觉层面包含大规模标注与精选困难干扰样本子集,用于鲁棒性评估;涵盖大量真实高速公路图像。多模态层面引入PaveVQA,一个支持单轮、多轮及专家修正交互的真实图像问答数据集,覆盖识别、定位、量化估计与养护推理。我们评估了多种前沿方法并提供详细分析,还提出一种融合领域模型作为工具的简单高效代理增强视觉问答框架。数据集已公开于:https://huggingface.co/datasets/MML-Group/PaveBench。

原文摘要 · Abstract (English)

Pavement condition assessment is essential for road safety and maintenance. Existing research has made significant progress. However, most studies focus on conventional computer vision tasks such as classification, detection, and segmentation. In real-world applications, pavement inspection requires more than visual recognition. It also requires quantitative analysis, explanation, and interactive decision support. Current datasets are limited. They focus on unimodal perception. They lack support for multi-turn interaction and fact-grounded reasoning. They also do not connect perception with vision-language analysis. To address these limitations, we introduce PaveBench, a large-scale benchmark for pavement distress perception and interactive vision-language analysis on real-world highway inspection images. PaveBench supports four core tasks: classification, object detection, semantic segmentation, and vision-language question answering. It provides unified task definitions and evaluation protocols. On the visual side, PaveBench provides large-scale annotations and includes a curated hard-distractor subset for robustness evaluation. It contains a large collection of real-world pavement images. On the multimodal side, we introduce PaveVQA, a real-image question answering (QA) dataset that supports single-turn, multi-turn, and expert-corrected interactions. It covers recognition, localization, quantitative estimation, and maintenance reasoning. We evaluate several state-of-the-art methods and provide a detailed analysis. We also present a simple and effective agent-augmented visual question answering framework that integrates domain-specific models as tools alongside vision-language models. The dataset is available at: https://huggingface.co/datasets/MML-Group/PaveBench.

道路检测视觉问答多模态智能运维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。