新式多模态大模型能精准定位遥感图像中的目标,适合零样本应用。
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
- 针对遥感图像设计具备细粒度空间推理能力的新模型
- 在零样本场景下实现高精度目标定位,尤其在低分辨率图像中表现优异
- 提供提示优化与失败案例分析,指导实际应用
多模态大语言模型(MLLM)在计算机视觉领域取得显著进展,尤其在零样本任务中表现突出。然而,其性能在地球观测(EO)图像等分布外场景中常出现下降。已有研究发现,传统MLLM在图像描述和场景理解上表现良好,但在需要精细空间推理的任务(如目标定位)中表现不佳。本文评估了近期专为提升空间推理能力而训练的MLLM,在遥感图像目标定位任务上的表现。结果表明,这些模型在特定条件下具有较强性能,尤其适用于零样本场景。同时,本文深入探讨了提示设计、地面采样距离(GSD)优化及失败案例分析,为后续评估与优化提供实用参考。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings. Unfortunately, their strong performance does not always transfer to out-of-distribution domains, such as earth observation (EO) imagery. Prior work has demonstrated that MLLMs excel at some EO tasks, such as image captioning and scene understanding, while failing at tasks that require more fine-grained spatial reasoning, such as object localization. However, MLLMs are advancing rapidly and insights quickly become out-dated. In this work, we analyze more recent MLLMs that have been explicitly trained to include fine-grained spatial reasoning capabilities, benchmarking them on EO object localization tasks. We demonstrate that these models are performant in certain settings, making them well suited for zero-shot scenarios. Additionally, we provide a detailed discussion focused on prompt selection, ground sample distance (GSD) optimization, and analyzing failure cases. We hope that this work will prove valuable as others evaluate whether an MLLM is well suited for a given EO localization task and how to optimize it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。