arXiv:2607.20284cs.CV2026-07

对比专用与通用模型,发现通用大模型在遥感理解上表现不逊于专业模型。

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

论文配图:Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
图 1 · 摘自论文原文
  • 系统评估遥感多模态大模型性能边界与跨任务泛化能力
  • 通用视觉大模型在无遥感微调下已能超越或媲美专用模型
  • 揭示空间推理、细粒度理解等关键短板,指导未来研究方向

多模态大语言模型(MLLMs)为遥感图像场景理解(RSISU)提供了自然语言交互的新范式。然而,现有遥感专用多模态大模型(RS-MLLMs)的能力边界、跨任务泛化性及任务特异性局限仍缺乏系统认识。本文对RSISU领域的MLLMs进行系统综述与诊断评估,回顾其技术演进,涵盖模型设计、多模态学习、训练数据与下游能力。进一步在多样化的遥感任务与基准测试中,将RS-MLLMs与通用计算机视觉多模态大模型(CV-MLLMs)进行对比。结果显示,尽管RS-MLLMs在特定任务中保持竞争力,尤其在遥感视觉定位与高分辨率视觉问答任务中;但更值得注意的是,通用CV-MLLMs在无需遥感领域微调的情况下,已在多项遥感任务中达到甚至超过专用模型的表现。这表明通用模型具有强迁移能力,当前的专用模型并未在所有任务中持续领先。此外,现有模型在空间与关系推理、细粒度视觉理解、指令多样性以及异构任务格式泛化方面仍存在明显局限。基于此,本文提出未来研究方向:可靠评估体系、多模态与高分辨率推理、高效部署,以及工具增强型遥感智能体。本综述为构建鲁棒、可泛化且实用的遥感多模态大模型提供系统参考。

原文摘要 · Abstract (English)

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.

遥感理解多模态大模型通用模型系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。