arXiv:2503.11807cs.CVcs.AI2025-03中稿 · IEEE India Geoscie…

用多时相哨兵2号数据清洗作物分类错误标签,提升模型准确率。

Mitigating Bad Ground Truth in Supervised Machine Learning based Crop Classification: A Multi-Level Framework with Sentinel-2 Images

  • 通过农田嵌入聚类识别相似作物特征,定位异常样本
  • 清理后随机森林模型F1分数提升最高达70个百分点
  • 适合农业遥感、精准农业决策等场景使用

在农业管理中,精确的地面真值(GT)数据对机器学习驱动的作物分类至关重要。然而,作物误标和地块标识错误普遍存在。本文提出一种多层级GT清洗框架,结合多时相哨兵2号(Sentinel-2)数据,通过生成农田嵌入、聚类相似作物模式并识别异常点来检测GT错误。利用假彩色合成图(FCC)验证聚类结果,并采用距离度量实现验证过程的自动化与可扩展。对比训练于清洗前后数据的模型发现:以干净GT数据训练的随机森林模型,其F1分数最高提升达70个百分点。该方法显著改进作物分类技术,有望应用于信贷评估与农业决策支持。

原文摘要 · Abstract (English)

In agricultural management, precise Ground Truth (GT) data is crucial for accurate Machine Learning (ML) based crop classification. Yet, issues like crop mislabeling and incorrect land identification are common. We propose a multi-level GT cleaning framework while utilizing multi-temporal Sentinel-2 data to address these issues. Specifically, this framework utilizes generating embeddings for farmland, clustering similar crop profiles, and identification of outliers indicating GT errors. We validated clusters with False Colour Composite (FCC) checks and used distance-based metrics to scale and automate this verification process. The importance of cleaning the GT data became apparent when the models were trained on the clean and unclean data. For instance, when we trained a Random Forest model with the clean GT data, we achieved upto 70\% absolute percentage points higher for the F1 score metric. This approach advances crop classification methodologies, with potential for applications towards improving loan underwriting and agricultural decision-making.

作物分类遥感数据清洗哨兵2号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。