测试三大遥感模型在农业任务中的跨区域迁移能力,发现其性能大幅下降。
Benchmarking Geospatial Foundation Models for Agriculture Applications

- 按地理区域划分训练验证测试集,评估模型跨区域泛化能力
- 三模型在新地区表现骤降,仅识别常见作物,忽略稀有作物
- 统一输入格式对不同模型影响各异,阻碍直接对比
预训练于卫星影像的地理空间基础模型有望在遥感任务和区域间实现广泛泛化,但其地理迁移能力尚未系统评估,尤其在农业应用中。本文构建了一个受控基准,评估 Prithvi、SpectralGPT 和 SatMAE 三个模型在美国内布拉斯加州、北卡罗来纳州、加利福尼亚州和明尼苏达州的多时相作物分割与变化检测任务中的表现。通过将训练、验证和测试集分别分配至不同区域,衡量模型在未见土地上的迁移效果。结果表明,三者均在区域分布偏移下显著退化,仅能预测最常见的作物,遗漏稀有作物。此外,将模型适配至统一输入格式后,各模型表现差异明显,加剧了架构间的直接比较难度。研究揭示了当前地理空间基础模型在农业应用中的关键局限,并提出区域感知评估应成为必要标准。
原文摘要 · Abstract (English)
Geospatial foundation models pretrained on satellite imagery promise broad generalization across remote sensing tasks and regions, but their geographic transferability has not been systematically tested, especially in agriculture applications. This paper presents a controlled benchmark that evaluates three models, Prithvi, SpectralGPT, and SatMAE, on multi-temporal crop segmentation and change detection across four U.S. states, Iowa, North Carolina, California, and Minnesota. By assigning each train, validation, and test split to a separate region, we measure how well each model transfers to land it has not seen. All three degrade sharply under regional distribution shift, predicting only the most common crops while missing rare ones. We further find that fitting these models to a shared input format affects each one differently, which complicates direct architectural comparison. These results expose key limitations of current geospatial foundation models for agriculture and point to region aware evaluation as a necessary standard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。