arXiv:2510.07791cs.CV2025-10被引 1

测试视觉语言模型在多摄像头地理时空推理能力,发现现有模型表现远低于人类。

GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models

  • 构建多视角地图与视频联动的地理时空推理评估基准
  • 最佳模型仅达34.9%准确率,远低于人类78.61%水平
  • 揭示模型对时空信息利用不均、预测能力弱、地图视频融合差等缺陷

近期,视觉-语言模型(VLMs)的时空智能受到广泛关注,因其在自动驾驶、具身智能和通用AI中的关键作用。现有时空评测基准主要聚焦于第一人称视角的图像/视频推理,或基于地图等图形上下文的地理推理,未能评估同时融合图像/视频与图形上下文的地理时空智能,而这在交通管理、应急响应等真实场景中至关重要。为此,我们提出地理时空推理基准(GTR-Bench),一个面向大规模摄像头网络中移动目标的地理时空推理新挑战。该基准更具挑战性,要求在地图与视频间多次切换视角,跨非重叠视域的多视频联合推理,并对未被任何视频覆盖的时空区域进行推断。对10余种主流VLMs的评估显示,即使最优专有模型Gemini-2.5-Pro(34.9%)也显著落后于人类表现(78.61%)。全面分析揭示当前模型在地理时空推理中的三大缺陷:(1) 空间与时间上下文利用失衡;(2) 时间预测能力弱,导致时序任务表现差;(3) 难以有效对齐与整合地图数据与多视角视频输入。我们认为GTR-Bench为时空智能研究与应用提供了重要洞见。基准与代码将开源至https://github.com/X-Luffy/GTR-Bench。

原文摘要 · Abstract (English)

Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Existing spatial-temporal benchmarks mainly focus on egocentric (first-person) perspective reasoning using images/video contexts, or geographic reasoning with graphical context (e.g., maps), thus fail to assess VLMs' geographic spatial-temporal intelligence that requires integrating both images/video and graphical context, which is crucial for real-world scenarios such as traffic management and emergency response. To address the gaps, we introduce Geo-Temporal Reasoning benchmark (GTR-Bench), a novel challenge for geographic temporal reasoning of moving targets in a large-scale camera network. GTR-Bench is more challenging as it requires multiple perspective switches between maps and videos, joint reasoning across multiple videos with non-overlapping fields of view, and inference over spatial-temporal regions that are unobserved by any video context. Evaluations of more than 10 popular VLMs on GTR-Bench show that even the best proprietary model, Gemini-2.5-Pro (34.9\%), significantly lags behind human performance (78.61\%) on geo-temporal reasoning. Moreover, our comprehensive analysis on GTR-Bench reveals three major deficiencies of current models for geo-temporal reasoning. (1) VLMs exhibit imbalanced utilization of spatial and temporal context during reasoning. (2) they show weak temporal forecasting ability, leading to poorer performance on temporally focused tasks. (3) they lack the capability to effectively align and integrate map data with multi-view video inputs. We believe GTR-Bench offers valuable insights and opens up new opportunities for research and applications in spatial-temporal intelligence. Benchmark and code will be released at https://github.com/X-Luffy/GTR-Bench.

视觉语言模型时空推理地理智能多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。