构建遥感多模态基准,揭示现有模型在空间理解上的短板
A Vision Centric Remote Sensing Benchmark
- 设计RSMMVP基准,识别CLIP模型误判的遥感图像对
- 实测显示主流模型在遥感视觉定位上表现严重不足
- 适合遥感智能解译与多模态模型改进的研究者参考
多模态大语言模型(MLLMs)在视觉-语言任务中取得显著进展,但其在遥感(RS)领域的应用仍相对滞后。与自然图像不同,遥感图像在视觉定位和空间推理方面具有独特挑战,当前的基于CLIP的MLLMs难以有效应对,尤其在区分视觉差异明显但语义相似的遥感图像时表现不佳。为此,本文提出遥感多模态视觉模式基准(RSMMVP),用于评估MLLMs在遥感任务中的能力,重点识别CLIP盲区样本——即视觉差异大但被错误赋予高相似度的遥感图像对。通过视觉问答(VQA)评测,分析了先进MLLMs在遥感场景下的性能,发现其在遥感特定表征学习方面存在显著局限。结果揭示了基于CLIP的视觉编码器在遥感任务中的根本性缺陷,并为未来开发更适用于遥感应用的高效多模态模型提供了基础。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike natural images, RS imagery presents unique challenges that current MLLMs struggle to handle, particularly in visual grounding and spatial reasoning. This study investigates the limitations of CLIP-based MLLMs in RS, highlighting their failure to differentiate visually distinct yet semantically similar RS images. To address this, we introduce a remote sensing multimodal visual patterns (RSMMVP) benchmark. It is designed to evaluate MLLMs in RS tasks by identifying the CLIP-blind pairs, where CLIP-based models incorrectly assign high similarity scores to visually distinct RS images. Through a visual question answering (VQA) evaluation, we analyze the performance of state-of-the-art MLLMs, revealing significant limitations in RS specific representation learning. The results provide valuable insights into the weaknesses of CLIP-based visual encoding and offer a foundation for future research to develop more effective MLLMs tailored for remote sensing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。