arXiv:2509.04334cs.CV2025-09ACL被引 3

用真人偏好评估大模型地理推理能力,更真实可靠。

GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models

  • 构建动态对比框架,让人类评价模型推理过程而非仅看答案
  • 基于数千次人工判断,评测17个前沿视觉语言模型表现
  • 适合关注地理推理、可解释AI的研究者与开发者

地理推理是需要结合视觉证据与空间常识推断合理位置的核心认知能力。尽管大视觉语言模型(LVLMs)取得进展,现有评估仍以结果为中心,依赖静态数据集和预定义标签,与开放世界推理任务不匹配。此类评估往往只关注标签匹配,忽视推理链条的合理性。本文提出GeoArena,一个基于人类偏好的动态评估框架,将地理推理评测转化为对真实图像中模型解释的成对质量比较,由人类评委从推理质量、证据融合和合理性三方面打分。我们公开部署该平台,对17个前沿LVLM进行了大规模人类评判,补充了现有基准,并推动具备地理知识与人类对齐的AI系统发展。进一步分析揭示了人类偏好的一致性及影响判断的关键因素。

原文摘要 · Abstract (English)

Geographic reasoning is a fundamental cognitive capability that requires models to infer plausible locations by synthesizing visual evidence with spatial world knowledge. Despite recent advances in large vision-language models (LVLMs), existing evaluation paradigms remain largely outcome-centric, relying on static datasets and predefined labels that are conceptually misaligned with open-world geographic inference. Such outcome-centric evaluations often focus exclusively on label matching, leaving the underlying linguistic reasoning chains as unexamined black boxes. In this work, we introduce GeoArena, a dynamic, human-preference-based evaluation framework for benchmarking open-world geographic reasoning. GeoArena reframes evaluation as a pairwise reasoning alignment task on in-the-wild images, where human judges compare model-generated explanations based on reasoning quality, evidence synthesis, and plausibility. We deploy GeoArena as a public platform and benchmark 17 frontier LVLMs using thousands of human judgments, which complements existing benchmarks and supports the development of geographically grounded, human-aligned AI systems. We further provide detailed analyses of model behavior, including reliability of human preferences and factors influencing judgments of geographic reasoning quality.

地理推理视觉语言模型人类偏好评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。