SatVision-TOA用海量云天遥感数据训练,提升气象监测精度。
SatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing Imagery
- 基于掩码图像建模自监督训练,无需标签
- 30亿参数模型在1亿张影像上预训练,mIOU达0.46
- 适合大气、云层与地表监测,尤其擅长粗分辨率全天空数据
基础模型有望通过大规模遥感数据预训练,推动遥感分析变革。现有模型多针对高分辨率无云影像,难以应对频繁时序监测或宽谱遥感需求。本文提出SatVision-TOA,首个基于14波段MODIS L1B地表顶辐射率(TOA)影像的遥感基础模型,采用SwinV2架构与掩码图像建模(MIM)框架,在1亿张影像上进行自监督预训练,参数量达30亿,是目前最大纯遥感影像预训练模型。实验表明,该模型在3D云检测等下游任务中表现优异,平均交并比(mIOU)达0.46,显著优于基线(0.22),且误检率降低超50%。模型通过学习多样大气与气溶胶条件,有效提升云与地表监测能力。
原文摘要 · Abstract (English)
Foundation models have the potential to transform the landscape of remote sensing (RS) data analysis by enabling large computer vision models to be pre-trained on vast amounts of remote sensing data. These models can then be fine-tuned with small amounts of labeled training and applied to a variety of applications. Most existing foundation models are designed for high spatial resolution, cloud-free satellite imagery or photos, limiting their applicability in scenarios that require frequent temporal monitoring or broad spectral profiles. As a result, foundation models trained solely on cloud-free images have limited utility for applications that involve atmospheric variables or require atmospheric corrections. We introduce SatVision-TOA, a novel foundation model pre-trained on 14-band MODIS L1B Top-Of-Atmosphere (TOA) radiance imagery, addressing the need for models pre-trained to handle moderate- and coarse-resolution all-sky remote sensing data. The SatVision-TOA model is pre-trained using a Masked-Image-Modeling (MIM) framework and the SwinV2 architecture, and learns detailed contextual representations through self-supervised learning without the need for labels. It is a 3 billion parameter model that is trained on 100 million images. To our knowledge this is the largest foundation model trained solely on satellite RS imagery. Results show that SatVision-TOA achieves superior performance over baseline methods on downstream tasks such as 3D cloud retrieval. Notably, the model achieves a mean intersection over union (mIOU) of 0.46, a substantial improvement over the baseline mIOU of 0.22. Additionally, the rate of false negative results in the fine-tuning task were reduced by over 50% compared to the baseline. Our work advances pre-trained vision modeling for multispectral RS by learning from a variety of atmospheric and aerosol conditions to improve cloud and land surface monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。