用多源时序遥感数据训练农业专用大模型,提升作物制图精度。
AgriFM: A Multi-source Temporal Remote Sensing Foundation Model for Agriculture Mapping
- 设计同步时空下采样的视频Swin结构,统一处理长序列遥感数据。
- 基于2500万+样本预训练,融合MODIS、Landsat和Sentinel-2三源数据。
- 支持多种下游任务,优于现有通用遥感模型和深度学习方法。
精准作物制图依赖于对多尺度时空模式的建模,空间尺度涵盖单个田块纹理到景观级上下文,时间尺度则包含短期物候变化与完整生长季动态。基于Transformer的遥感基础模型(RSFMs)因具备统一时空处理能力而展现出巨大潜力,但现有模型仍不理想:或使用固定时空窗口忽略作物系统的多尺度特性,或完全忽略时间信息而仅关注空间模式。为此,我们提出AgriFM——专为农业作物制图设计的多源遥感基础模型。通过验证同步分层时空特征提取的必要性,我们改进了Video Swin Transformer架构,使时间下采样与空间缩放同步进行,从而高效统一处理长时间序列卫星数据。AgriFM融合来自MODIS、Landsat-8/9和Sentinel-2的时序丰富数据流,并在包含超过2500万图像样本的全球代表性数据集上进行预训练,标注依据为土地覆盖产品。最终框架采用可动态融合多源时空表征的灵活解码器,支持多样化下游任务。全面评估表明,AgriFM在所有下游任务中均显著优于传统深度学习方法及当前最先进的通用遥感基础模型。
原文摘要 · Abstract (English)
Accurate crop mapping fundamentally relies on modeling multi-scale spatiotemporal patterns, where spatial scales range from individual field textures to landscape-level context, and temporal scales capture both short-term phenological transitions and full growing-season dynamics. Transformer-based remote sensing foundation models (RSFMs) offer promising potential for crop mapping due to their innate ability for unified spatiotemporal processing. However, current RSFMs remain suboptimal for crop mapping: they either employ fixed spatiotemporal windows that ignore the multi-scale nature of crop systems or completely disregard temporal information by focusing solely on spatial patterns. To bridge these gaps, we present AgriFM, a multi-source remote sensing foundation model specifically designed for agricultural crop mapping. Our approach begins by establishing the necessity of simultaneous hierarchical spatiotemporal feature extraction, leading to the development of a modified Video Swin Transformer architecture where temporal down-sampling is synchronized with spatial scaling operations. This modified backbone enables efficient unified processing of long time-series satellite inputs. AgriFM leverages temporally rich data streams from three satellite sources including MODIS, Landsat-8/9 and Sentinel-2, and is pre-trained on a global representative dataset comprising over 25 million image samples supervised by land cover products. The resulting framework incorporates a versatile decoder architecture that dynamically fuses these learned spatiotemporal representations, supporting diverse downstream tasks. Comprehensive evaluations demonstrate AgriFM's superior performance over conventional deep learning approaches and state-of-the-art general-purpose RSFMs across all downstream tasks. Codes will be available at https://github.com/flyakon/AgriFM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。