arXiv:2411.18145cs.CV2024-11NeurIPS被引 19

构建首个面向遥感的视觉语言模型评测基准,全面评估感知与推理能力。

CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models

  • 设计六维度23子任务,覆盖遥感核心能力
  • 基于50城数据构建10507道客观题,保证评测质量
  • 揭示主流模型在遥感场景下的显著能力短板

大型视觉语言模型(VLMs)在地球观测任务中展现出卓越的感知与推理能力,但针对其遥感能力的系统性评测仍属空白。为此,我们提出CHOICE,一个全面的遥感能力评测基准,聚焦感知与推理两大核心维度,进一步细分为六个二级维度和二十三个具体任务,确保评估覆盖全面。通过从全球50座城市收集数据、严格设计问题并进行质量控制,确保全部10,507个问题的高质量。新构建的数据集采用多选题形式并配有明确答案,支持客观、直接的性能评估。我们对3个专有模型和21个开源模型的评估揭示了它们在特定遥感场景中的关键局限。我们希望CHOICE能成为该领域的重要资源,推动对VLMs在遥感中挑战与潜力的深入理解。相关代码与数据将公开于[此链接](https://github.com/ShawnAn-WHU/CHOICE)。

原文摘要 · Abstract (English)

The rapid advancement of Large Vision-Language Models (VLMs), both general-domain models and those specifically tailored for remote sensing, has demonstrated exceptional perception and reasoning capabilities in Earth observation tasks. However, a benchmark for systematically evaluating their capabilities in this domain is still lacking. To bridge this gap, we propose CHOICE, an extensive benchmark designed to objectively evaluate the hierarchical remote sensing capabilities of VLMs. Focusing on 2 primary capability dimensions essential to remote sensing: perception and reasoning, we further categorize 6 secondary dimensions and 23 leaf tasks to ensure a well-rounded assessment coverage. CHOICE guarantees the quality of all 10,507 problems through a rigorous process of data collection from 50 globally distributed cities, question construction and quality control. The newly curated data and the format of multiple-choice questions with definitive answers allow for an objective and straightforward performance assessment. Our evaluation of 3 proprietary and 21 open-source VLMs highlights their critical limitations within this specialized context. We hope that CHOICE will serve as a valuable resource and offer deeper insights into the challenges and potential of VLMs in the field of remote sensing. We will release CHOICE at [this https URL](https://github.com/ShawnAn-WHU/CHOICE).

遥感视觉语言模型评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。