arXiv:2607.27278cs.CV2026-07

构建首个覆盖广类别与多查询形式的遥感开放词汇评估基准

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

论文配图:OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
图 1 · 摘自论文原文
  • 设计涵盖层级类别与正负表达的宽泛类别体系
  • 验证当前方法性能有限,大模型表现最优
  • 适合关注遥感语义理解与评估基准研究者

开放词汇遥感(EO)旨在定位自然语言描述的地理空间概念,而非固定标签集。现有基准通常覆盖类别窄或查询形式单一。为此,我们提出 OVEarth-Bench,从两个维度扩展评估:类别广度(通过包含正负表达的多层次类别覆盖)和查询多样性(涵盖词汇、指代与推理类查询)。该基准支持掩码与边界框定位,采用统一零样本协议。我们评估了多种通用与遥感专用方法。结果表明:(1) 当前方法性能仍受限,但更广类别覆盖使模型排名更稳定;(2) 基于多模态大模型的方法整体表现最强;(3) 遥感专用方法普遍弱于通用模型,极少能超越顶尖方法。这些发现为未来开放词汇遥感方法设计提供指导,并强调构建更真实、多样、高质量、大规模基准的重要性。数据与评估工具包已公开于 https://earth-insights.github.io/OVEarth-bench。

原文摘要 · Abstract (English)

Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with positive and negative expressions, and query diversity, through vocabulary, referring, and reasoning queries. The benchmark supports mask and box localization under a unified zero-shot protocol. We evaluate a broad set of general and EO-specific methods. The evaluation reveals that: (1) the performance of current methods remains limited, while broader category coverage yields more stable model rankings; (2) MLLM-based methods achieve the strongest overall performance; and (3) EO-specific methods generally underperform general models and rarely match the strongest methods. These findings provide guidance for future open-vocabulary EO method design and highlight the importance of developing more realistic, diverse, high-quality, and large-scale benchmarks for reliable evaluation. Our data and evaluation package are released at https://earth-insights.github.io/OVEarth-bench.

遥感开放词汇评估基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。