构建城市多视角空间推理基准,评估视觉语言模型在复杂城市环境中的理解能力。
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
- 设计四类视角动态模拟真实拍摄变化,覆盖车、无人机、卫星等多平台视角。
- 包含5022个标注精细的多视图问答对,准确率最高仅54.1%,远低于人类水平。
- 揭示大模型与人类在空间推理上的根本差距,适合评估城市场景下的认知能力。
跨视角空间推理对具身智能至关重要,支撑复杂环境中的空间理解、心理模拟与规划。现有基准多聚焦室内或街景,忽视开放城市空间中丰富的语义、复杂几何与视角差异带来的挑战。为此,我们提出CityCube,一个系统性基准,用于评估当前视觉语言模型(VLMs)在城市环境中的跨视角推理能力。CityCube整合四类视角动态以模拟摄像机运动,涵盖车辆、无人机、卫星等多平台的广泛视角。为全面评估,其包含5,022个精心标注的多视图问答对,按五种认知维度和三种空间关系表达分类。对33个VLM的综合评估显示显著性能差距:即使大型模型最高准确率也未超过54.1%,比人类表现低34.2%。相反,小规模微调后的VLM可达到60.0%以上准确率,凸显本基准的重要性。进一步分析表明任务间存在相关性,且VLM与人类推理存在根本认知差异。
原文摘要 · Abstract (English)
Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily emphasize indoor or street settings, overlooking the unique challenges of open-ended urban spaces characterized by rich semantics, complex geometries, and view variations. To address this, we introduce CityCube, a systematic benchmark designed to probe cross-view reasoning capabilities of current VLMs in urban settings. CityCube integrates four viewpoint dynamics to mimic camera movements and spans a wide spectrum of perspectives from multiple platforms, e.g., vehicles, drones and satellites. For a comprehensive assessment, it features 5,022 meticulously annotated multi-view QA pairs categorized into five cognitive dimensions and three spatial relation expressions. A comprehensive evaluation of 33 VLMs reveals a significant performance disparity with humans: even large-scale models struggle to exceed 54.1% accuracy, remaining 34.2% below human performance. By contrast, small-scale fine-tuned VLMs achieve over 60.0% accuracy, highlighting the necessity of our benchmark. Further analyses indicate the task correlations and fundamental cognitive disparity between VLMs and human-like reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。