arXiv:2507.00356cs.CVcs.AI2025-07被引 1

基于吉林一号卫星的高分辨率遥感大模型,提升影像智能解译能力

CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation

  • 针对吉林一号卫星设计多尺度模型,总参数达21亿
  • 构建全球覆盖1500万样本自监督数据集,支持四季对比学习
  • 在10个基准数据集上达到领先效果,适合高精度遥感应用

深度学习显著推动了遥感智能解译的发展,基于大规模预训练的视觉基础模型正在重塑地球观测领域。然而,与中分辨率数据的开放获取和高时空覆盖相比,超高清光学遥感影像获取渠道有限,制约了高分辨率遥感视觉基础模型(RSVFM)的发展。作为全球最大的亚米级商业遥感卫星星座,吉林一号拥有丰富的亚米级影像资源。本文提出CGEarthEye,一个专为吉林一号卫星特性设计的RSVFM框架,包含五个不同参数规模的骨干网络,总参数量达21亿。为增强模型表征能力,我们构建了首个1500万级、全球覆盖、按季度采样的多时相自监督学习(SSL)数据集JLSSD,通过多层次表示聚类与采样策略实现。框架融合季节对比、基于增强的对比和掩码块令牌对比策略进行预训练。在涵盖四类典型遥感任务的10个基准数据集上全面评估显示,CGEarthEye持续达到最先进性能。进一步分析表明,该模型在特征可视化、收敛性、参数效率及实际制图应用方面表现优异。本研究预期CGEarthEye卓越的表征能力将推动吉林一号数据在传统地球观测应用中的更广泛高效使用。

原文摘要 · Abstract (English)

Deep learning methods have significantly advanced the development of intelligent rinterpretation in remote sensing (RS), with foundational model research based on large-scale pre-training paradigms rapidly reshaping various domains of Earth Observation (EO). However, compared to the open accessibility and high spatiotemporal coverage of medium-resolution data, the limited acquisition channels for ultra-high-resolution optical RS imagery have constrained the progress of high-resolution remote sensing vision foundation models (RSVFM). As the world's largest sub-meter-level commercial RS satellite constellation, the Jilin-1 constellation possesses abundant sub-meter-level image resources. This study proposes CGEarthEye, a RSVFM framework specifically designed for Jilin-1 satellite characteristics, comprising five backbones with different parameter scales with totaling 2.1 billion parameters. To enhance the representational capacity of the foundation model, we developed JLSSD, the first 15-million-scale multi-temporal self-supervised learning (SSL) dataset featuring global coverage with quarterly temporal sampling within a single year, constructed through multi-level representation clustering and sampling strategies. The framework integrates seasonal contrast, augmentation-based contrast, and masked patch token contrastive strategies for pre-training. Comprehensive evaluations across 10 benchmark datasets covering four typical RS tasks demonstrate that the CGEarthEye consistently achieves state-of-the-art (SOTA) performance. Further analysis reveals CGEarthEye's superior characteristics in feature visualization, model convergence, parameter efficiency, and practical mapping applications. This study anticipates that the exceptional representation capabilities of CGEarthEye will facilitate broader and more efficient applications of Jilin-1 data in traditional EO application.

遥感视觉基础模型吉林一号自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。