构建首个面向无人机道路损毁定位的视觉语言基准,测试模型与智能体在真实场景下的定位能力。
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

- 将视觉语言模型与大模型驱动智能体结合,统一评估图像与提示的定位性能。
- 闭源模型领先但仍有超50%性能缺口,开源模型表现更差且小目标定位失效。
- 适用于研究视觉定位、自主智能体与真实世界泛化能力的学者与工程师。
我们提出WildRoadBench,一个耦合视觉语言模型直接视觉定位与大模型驱动智能体自主科研工程的野生航拍道路损毁定位基准,基于同一专业标注的无人机数据集。采用相同图像集与每类AP_50指标,在两种协议下评估:VLM赛道测试固定模型在单图单提示下,通过统一提示、解码与解析流程定位特定损毁的能力;Agent赛道测试智能体仅凭任务简述、少量探索样本和固定交互预算,能否搜索网络、适配预训练组件、编写训练与推理代码,并通过标量反馈接口提交预测。我们对大量闭源前沿模型、开源视觉语言模型及多个前沿大模型驱动智能体进行基准测试。两者在真实场景中均未达可靠性能:闭源模型虽领先,但仍遗留超50%指标未达成;开源模型整体表现滞后,新版本或推理型变体未持续提升;小目标定位对所有开源模型均失效;智能体虽具更强能力仍落后最强VLM,部分未能于预算内提交有效结果。代码与数据已公开,支持可复现研究。
原文摘要 · Abstract (English)
We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus. The same image set and the same per-class AP_50 metric are evaluated under two protocols. The VLM Track measures whether a fixed VLM can localise domain-specific damage from one image and one short prompt under a unified prompting, decoding and parsing pipeline. The Agent Track measures whether an autonomous agent, given only a written task brief, a small exploratory slice and a fixed interaction budget, can search the public web, adapt pretrained components, write training and inference code, and submit predictions through a scalar-feedback oracle on a hidden holdout. We benchmark a broad pool of closed-source frontier models and open-source VLMs together with several frontier LLM-driven agents. Both routes remain far from reliable performance in this wild setting: closed-source frontier models lead the VLM leaderboard but still leave more than half of the metric on the table; open-source grounders plateau well below them, and newer generations or reasoning-style variants do not consistently improve grounding; small targets collapse for every open-source model; agents lag the strongest VLM despite richer affordances, and several fail to land a valid submission within the budget. We release the code and data at https://anonymous.4open.science/r/wildroadbench-0607 to support reproducible follow-up research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。