arXiv:2507.20776cs.CV2025-07TPAMI被引 20

RingMo-Agent统一处理多源遥感数据,支持跨平台跨模态推理。

RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning

  • 基于分离嵌入层构建异构模态特征,减少跨模态干扰。
  • 在300万+图文对上训练,覆盖光学、SAR、红外及卫星、无人机数据。
  • 支持用户指令驱动的感知与复杂分析任务,通用性强。

多源遥感图像因传感器特性与成像视角差异,呈现出丰富的细节。现有遥感视觉-语言研究多依赖相对同质的数据源,且局限于分类、描述等传统视觉任务,难以作为独立统一框架应对真实场景中多样遥感数据。为此,我们提出RingMo-Agent,一个面向多模态、多平台的统一遥感基础模型,可依据用户文本指令完成感知与推理任务。相比现有模型,该模型:1)依托大规模遥感视觉-语言数据集RS-VL3M,包含超300万张图像-文本对,覆盖光学、合成孔径雷达(SAR)、红外(IR)模态及卫星、无人机平台,涵盖感知与复杂推理任务;2)通过分离嵌入层学习模态自适应表示,为异构模态构建独立特征,降低跨模态干扰;3)通过任务特定令牌与基于令牌的高维隐状态解码机制,统一建模各类任务,尤其适用于长时序空间分析。在多种遥感视觉-语言任务上的实验表明,RingMo-Agent不仅在视觉理解与复杂分析任务中表现优异,且在不同平台与传感模态间具备强泛化能力。

原文摘要 · Abstract (English)

Remote sensing (RS) images from multiple modalities and platforms exhibit diverse details due to differences in sensor characteristics and imaging perspectives. Existing vision-language research in RS largely relies on relatively homogeneous data sources. Moreover, they still remain limited to conventional visual perception tasks such as classification or captioning. As a result, these methods fail to serve as a unified and standalone framework capable of effectively handling RS imagery from diverse sources in real-world applications. To address these issues, we propose RingMo-Agent, a model designed to handle multi-modal and multi-platform data that performs perception and reasoning tasks based on user textual instructions. Compared with existing models, RingMo-Agent 1) is supported by a large-scale vision-language dataset named RS-VL3M, comprising over 3 million image-text pairs, spanning optical, SAR, and infrared (IR) modalities collected from both satellite and UAV platforms, covering perception and challenging reasoning tasks; 2) learns modality adaptive representations by incorporating separated embedding layers to construct isolated features for heterogeneous modalities and reduce cross-modal interference; 3) unifies task modeling by introducing task-specific tokens and employing a token-based high-dimensional hidden state decoding mechanism designed for long-horizon spatial tasks. Extensive experiments on various RS vision-language tasks demonstrate that RingMo-Agent not only proves effective in both visual understanding and sophisticated analytical tasks, but also exhibits strong generalizability across different platforms and sensing modalities.

遥感多模态基础模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。