arXiv:2504.03166cs.CV2025-04TPAMI被引 29

RingMoE统一多模态遥感模型,147亿参数,提升图像解析准确率

RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation

  • 分层专家混合架构,融合光学、SAR等多源数据
  • 在23项任务中刷新最优性能,支持单/多模态切换
  • 动态剪枝可压缩至10亿参数,适合实际部署

基础模型的快速发展推动了视觉表征学习的自监督范式,但在遥感领域仍受限于单一或有限模态处理能力。光学、合成孔径雷达(SAR)和多光谱数据具有互补性,能显著降低单源分析中的模糊性和不确定性。为此,我们提出RingMoE,一个147亿参数的统一多模态遥感基础模型,基于9颗卫星的4亿张多模态遥感图像进行预训练。其三大创新包括:(1) 分层专家混合(MoE)架构,包含模态专用、协作和共享专家,有效建模模态内知识并捕捉跨模态依赖,缓解表示冲突;(2) 物理感知自监督学习,将传感器特有辐射特性嵌入预训练目标;(3) 动态专家剪枝,实现从147亿到10亿参数的自适应压缩,同时保持性能,便于地球观测场景部署。在涵盖分类、检测、分割、跟踪、变化检测和深度估计六类任务的23个基准上,RingMoE优于现有基础模型,达到新SOTA。该模型已应用于应急响应、土地管理、海洋科学与城市规划等多个领域。

原文摘要 · Abstract (English)

The rapid advancement of foundation models has revolutionized visual representation learning in a self-supervised manner. However, their application in remote sensing (RS) remains constrained by a fundamental gap: existing models predominantly handle single or limited modalities, overlooking the inherently multi-modal nature of RS observations. Optical, synthetic aperture radar (SAR), and multi-spectral data offer complementary insights that significantly reduce the inherent ambiguity and uncertainty in single-source analysis. To bridge this gap, we introduce RingMoE, a unified multi-modal RS foundation model with 14.7 billion parameters, pre-trained on 400 million multi-modal RS images from nine satellites. RingMoE incorporates three key innovations: (1) A hierarchical Mixture-of-Experts (MoE) architecture comprising modal-specialized, collaborative, and shared experts, effectively modeling intra-modal knowledge while capturing cross-modal dependencies to mitigate conflicts between modal representations; (2) Physics-informed self-supervised learning, explicitly embedding sensor-specific radiometric characteristics into the pre-training objectives; (3) Dynamic expert pruning, enabling adaptive model compression from 14.7B to 1B parameters while maintaining performance, facilitating efficient deployment in Earth observation applications. Evaluated across 23 benchmarks spanning six key RS tasks (i.e., classification, detection, segmentation, tracking, change detection, and depth estimation), RingMoE outperforms existing foundation models and sets new SOTAs, demonstrating remarkable adaptability from single-modal to multi-modal scenarios. Beyond theoretical progress, it has been deployed and trialed in multiple sectors, including emergency response, land management, marine sciences, and urban planning.

遥感多模态专家混合自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。