arXiv:2504.12795cs.CV2025-04被引 15

首个支持多源遥感影像多层级理解的视觉提示模型。

EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting

  • 引入文本+点/框/自由形状的双重提示机制,模拟人类指代行为。
  • 在多源遥感数据上实现从场景到细粒度物体属性的分层推理,性能显著提升。
  • 适合需要跨模态、多任务、高精度空间分析的研究者与应用开发者。

近年来,自然领域多模态大模型通过视觉与文本提示展现了出色的空间推理能力,但其直接迁移至遥感(RS)领域面临传感物理异质性、模态多样性和独特空间尺度的挑战。现有遥感多模态大模型主要局限于光学影像和简单语言交互,限制了灵活可扩展的实际应用。本文提出EarthGPT-X,首个具备灵活性的空间多模态大模型,统一多源遥感影像理解,在单一框架内支持粗粒度与细粒度视觉任务,且能响应多种视觉提示。不同于以往模型,EarthGPT-X创新性地:1)设计双提示机制,结合文本指令与点、框、自由形状等视觉提示,模拟人类日常指代多样性;2)构建涵盖多源、多层级的提示数据集,使模型从整体图像理解延伸至分层空间推理,包括场景级理解与细粒度对象属性及关系分析;3)采用跨域单阶段融合训练策略,实现模态与任务间的高效一致对齐。大量实验表明,EarthGPT-X显著优于先前自然与遥感多模态大模型,首次建立了一个可在遥感场景中实现多源、多任务、多层级解读的视觉提示框架。

原文摘要 · Abstract (English)

Recent advances in natural-domain multi-modal large language models (MLLMs) have demonstrated effective spatial reasoning through visual and textual prompting. However, their direct transfer to remote sensing (RS) is hindered by heterogeneous sensing physics, diverse modalities, and unique spatial scales. Existing RS MLLMs are mainly limited to optical imagery and plain language interaction, preventing flexible and scalable real-world applications. In this article, EarthGPT-X is proposed, the first flexible spatial MLLM that unifies multi-source RS imagery comprehension and accomplishes both coarse-grained and fine-grained visual tasks under diverse visual prompts in a single framework. Distinct from prior models, EarthGPT-X introduces: 1) a dual-prompt mechanism combining text instructions with various visual prompts (i.e., point, box, and free-form) to mimic the versatility of referring in human life; 2) a comprehensive multi-source multi-level prompting dataset, the model advances beyond holistic image understanding to support hierarchical spatial reasoning, including scene-level understanding and fine-grained object attributes and relational analysis; 3) a cross-domain one-stage fusion training strategy, enabling efficient and consistent alignment across modalities and tasks. Extensive experiments demonstrate that EarthGPT-X substantially outperforms prior nature and RS MLLMs, establishing the first framework capable of multi-source, multi-task, and multi-level interpretation using visual prompting in RS scenarios.

遥感理解多模态模型视觉提示空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。