构建道路标记细粒度理解基准,评估大模型在城市道路中的空间推理能力。
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- 以道路标线为切入点,设计鸟瞰与第一视角图像的八类任务评测体系。
- 涵盖3040个人工验证案例,基于中国五城真实交通数据,测试模型识别与推理能力。
- 揭示主流大模型在城市场景中细粒度空间理解严重不足,部分表现低于随机基线。
多模态大语言模型(MLLMs)在通用空间理解与推理方面表现强大,但在复杂城市场景下的细粒度空间理解能力尚未受到足够关注。为此,本文聚焦道路标线这一典型细粒度空间要素,因其在城市交通网络中的关键作用。围绕道路标线与城市交通系统,我们提出 extbf{RoadBench},一个系统性基准,通过鸟瞰图(BEV)与第一视角图(FPV)输入,全面评估MLLMs在细粒度空间理解与推理方面的能力。该基准包含8项任务,共3,040个严格人工验证的测试案例,源自2,137张唯一的BEV图像和721张唯一的FPV图像,数据采集自五个交通规范相对一致的中国城市。这些任务构建了从局部空间理解到全局推理的系统性评估框架,不仅检验模型的识别、联合理解与推理能力,还评估其融合图像信息与领域知识的能力。对20个主流MLLMs的评估表明,RoadBench对现有模型构成显著挑战,暴露出其在城市场景中细粒度空间理解与推理能力的明显缺陷,在某些任务中甚至低于简单规则或随机选择基线。这些发现连同RoadBench本身,将推动MLLMs空间理解能力的全面提升。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine-grained spatial understanding and reasoning capabilities in complex urban scenarios have not received significant attention in the fields of both research and industry. To fill this gap, we focus primarily on road markings as a typical example of fine-grained spatial elements under urban scenarios, given the essential role of the integrated road traffic network they form within cities. Around road markings and urban traffic systems, we propose \textbf{RoadBench}, a systematic benchmark that comprehensively evaluates MLLMs' fine-grained spatial understanding and reasoning capabilities using Bird's-Eye View (BEV) and First-Person View (FPV) image inputs. This benchmark comprises eight tasks consisting of 3,040 strictly manually verified test cases, constructed from 2,137 unique BEV images and 721 unique FPV images collected from five Chinese cities with relatively consistent traffic conventions. These tasks form a systematic evaluation framework that bridges understanding at local spatial scopes to global reasoning. They not only test MLLMs' capabilities in recognition, joint understanding, and reasoning but also assess their ability to integrate image information with domain knowledge. After evaluating 20 mainstream MLLMs, we confirm that RoadBench is a challenging benchmark for MLLMs while revealing significant shortcomings in existing MLLMs' fine-grained spatial understanding and reasoning capabilities within urban scenarios. In certain tasks, their performance even falls short of simple rule-based or random selection baselines. These findings, along with RoadBench itself, will contribute to the comprehensive advancement of spatial understanding capabilities for MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。