构建全球农业地块分割基准数据集,助力机器学习精准识别各地农田边界。
Fields of The World: A Machine Learning Benchmark Dataset For Global Agricultural Field Boundary Segmentation

- 基于24国多时相多光谱遥感影像,构建7万+样本的全球农田实例分割数据集
- 模型在未训练过的国家上表现更优,零样本与微调效果显著提升
- 适用于农业监测、环境评估等领域的全球尺度智能分析
作物地块边界是农业监测与评估的基础数据,但人工获取成本高。利用机器学习从遥感图像中自动提取地块边界,可实现全球尺度的数据供给。然而现有方法在地理覆盖、精度和泛化能力方面仍有不足,且缺乏涵盖全球农业多样性的真实标注数据。本文提出「全球农田」(Fields of The World, FTW)——一个跨四大洲24个国家的农业地块实例分割机器学习基准数据集。FTW包含70,462个样本,每条样本配有实例与语义分割掩码,以及多时相、多光谱的哨兵-2卫星影像。我们提供了基线模型结果,表明在FTW上预训练的模型,在未见国家上展现出更好的零样本与微调性能;此外,在埃塞俄比亚真实场景中,其零样本推理也表现出良好效果。
原文摘要 · Abstract (English)
Crop field boundaries are foundational datasets for agricultural monitoring and assessments but are expensive to collect manually. Machine learning (ML) methods for automatically extracting field boundaries from remotely sensed images could help realize the demand for these datasets at a global scale. However, current ML methods for field instance segmentation lack sufficient geographic coverage, accuracy, and generalization capabilities. Further, research on improving ML methods is restricted by the lack of labeled datasets representing the diversity of global agricultural fields. We present Fields of The World (FTW) -- a novel ML benchmark dataset for agricultural field instance segmentation spanning 24 countries on four continents (Europe, Africa, Asia, and South America). FTW is an order of magnitude larger than previous datasets with 70,462 samples, each containing instance and semantic segmentation masks paired with multi-date, multi-spectral Sentinel-2 satellite images. We provide results from baseline models for the new FTW benchmark, show that models trained on FTW have better zero-shot and fine-tuning performance in held-out countries than models that aren't pre-trained with diverse datasets, and show positive qualitative zero-shot results of FTW models in a real-world scenario -- running on Sentinel-2 scenes over Ethiopia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。