arXiv:2603.09471cs.CV2026-03KDD被引 4

首个面向遥感视觉语言模型的综合评估基准,覆盖感知、推理与鲁棒性三维度。

OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks

  • 构建多源遥感数据下的28项细粒度任务,支持选择题与开放式问答
  • 包含9275张高质量图像和44210条人工验证指令,覆盖多元地理场景
  • 采用盲测与语义一致性要求,降低语言偏见,检验模型真实视觉理解能力

视觉语言模型在通用任务中展现出强大的感知与推理能力,引发其在地球观测领域应用的兴趣。然而,缺乏系统性的遥感视觉语言模型(RSVLM)评估基准。为此,我们提出OmniEarth,一个面向真实地球观测场景的RSVLM评估基准。OmniEarth从感知、推理与鲁棒性三个能力维度组织任务,涵盖28项细粒度任务,支持多源传感数据与多样地理环境。任务形式包括多项选择题VQA和开放式VQA,后者包含纯文本输出(描述任务)、边界框输出(视觉定位任务)及掩码输出(分割任务)。为减少语言偏差并检验模型是否依赖视觉证据,OmniEarth采用盲测协议与五重语义一致性要求。基准包含9,275张经严格质控的图像,包括来自吉林一号(JL-1)的专有卫星影像,以及44,210条人工验证的指令。我们对基于对比学习的模型、通用闭源与开源视觉语言模型及专用RSVLM进行了系统评估,结果表明现有模型在复杂地理任务上仍表现不佳,揭示了遥感应用中的显著差距。OmniEarth已公开发布于https://huggingface.co/datasets/sjeeudd/OmniEarth。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated effective perception and reasoning capabilities on general-domain tasks, leading to growing interest in their application to Earth observation. However, a systematic benchmark for comprehensively evaluating remote sensing vision-language models (RSVLMs) remains lacking. To address this gap, we introduce OmniEarth, a benchmark for evaluating RSVLMs under realistic Earth observation scenarios. OmniEarth organizes tasks along three capability dimensions: perception, reasoning, and robustness. It defines 28 fine-grained tasks covering multi-source sensing data and diverse geospatial contexts. The benchmark supports two task formulations: multiple-choice VQA and open-ended VQA. The latter includes pure text outputs for captioning tasks, bounding box outputs for visual grounding tasks, and mask outputs for segmentation tasks. To reduce linguistic bias and examine whether model predictions rely on visual evidence, OmniEarth adopts a blind test protocol and a quintuple semantic consistency requirement. OmniEarth includes 9,275 carefully quality-controlled images, including proprietary satellite imagery from Jilin-1 (JL-1), along with 44,210 manually verified instructions. We conduct a systematic evaluation of contrastive learning-based models, general closed-source and open-source VLMs, as well as RSVLMs. Results show that existing VLMs still struggle with geospatially complex tasks, revealing clear gaps that need to be addressed for remote sensing applications. OmniEarth is publicly available at https://huggingface.co/datasets/sjeeudd/OmniEarth.

遥感视觉语言模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。