首个专为地理空间任务设计的视觉语言模型评测基准
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
- 构建覆盖1万+人工验证指令的地理空间多任务评测集
- 现有模型在关键任务上最高仅41.7%准确率,远低于人类水平
- 适合研究遥感理解、城市规划与灾害监测的AI开发者使用
尽管近年出现诸多通用视觉-语言模型(VLM)评测基准,但它们未能有效应对地理空间应用的独特挑战。通用评测集缺乏对时序变化检测、大规模目标计数、微小目标识别及遥感图像中实体关系理解等关键问题的设计。为此,我们提出GEOBench-VLM,一个专为地理空间任务设计的综合性评测基准,涵盖场景理解、目标计数、定位、细粒度分类、分割和时序分析等任务。该基准包含超过10,000条经人工验证的指令,覆盖多样化的视觉条件、物体类型与尺度。我们评估了多个前沿VLM模型的表现,结果表明,尽管现有模型展现潜力,但在地理空间特定任务上仍面临显著挑战。其中表现最佳的LLaVa-OneVision在多项选择题上仅达41.7%准确率,略高于随机猜测(约25%),约为GPT-4o的两倍。该基准已公开于https://github.com/The-AI-Alliance/GEO-Bench-VLM。
原文摘要 · Abstract (English)
While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, an essential component for applications such as environmental monitoring, urban planning, and disaster management. Key challenges in the geospatial domain include temporal change detection, large-scale object counting, tiny object detection, and understanding relationships between entities in remote sensing imagery. To bridge this gap, we present GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization, fine-grained categorization, segmentation, and temporal analysis. Our benchmark features over 10,000 manually verified instructions and spanning diverse visual conditions, object types, and scales. We evaluate several state-of-the-art VLMs to assess performance on geospatial-specific challenges. The results indicate that although existing VLMs demonstrate potential, they face challenges when dealing with geospatial-specific tasks, highlighting the room for further improvements. Notably, the best-performing LLaVa-OneVision achieves only 41.7% accuracy on MCQs, slightly more than GPT-4o, which is approximately double the random guess performance. Our benchmark is publicly available at https://github.com/The-AI-Alliance/GEO-Bench-VLM .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。