arXiv:2607.23908cs.CV2026-07中稿 · ECCV

用地理空间嵌入检测并清理农作物数据中的错误标签。

Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets

论文配图:Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets
图 1 · 摘自论文原文
  • 基于预训练遥感编码器的嵌入向量,按区域和作物类型比对异常样本。
  • 在合成数据中检测准确率最高达0.84(AUROC),真实数据中提升模型精度。
  • 方法可复现且适用于大规模遥感参考数据清洗,适合农业遥感研究者。

高质量参考数据仍是全球作物类型制图的关键瓶颈。如WorldCereal等系统从地块登记、国家数据库、实地调查和地图产品等多种异源数据中聚合标签,每类数据均存在偏差、覆盖盲区和未知标签噪声。简单全局规则无效,因作物物候与观测条件在区域和季节间差异显著。本研究聚焦一个实际问题:基于地理空间基础模型生成的嵌入是否可用于参考数据清洗?我们提出一种实用、区域感知的嵌入式异常(EBA)检测框架,利用预训练地球观测编码器的嵌入向量,对同地区同作物样本进行对比,标记异常点,并测试剔除或降权这些样本后模型性能是否提升。结果表明:在合成真值下,检测器将注入的标签错误集中度提高2.5–5倍于随机水平(检测AUROC达0.84);在真实数据中,独立模型验证显示,移除或置信加权被标记的保留样本能提升作物类型与土地覆盖模型的准确率。在五个宏观区域上,基于固定预留划分,应用标记结果改进了WorldCereal作物类型模型。保守清洗有益,过度清洗则有害。EBA检测方法设计为可复现、可扩展,可作为清洁大型、嘈杂地球观测参考数据集的模板,超越作物映射范畴。

原文摘要 · Abstract (English)

High quality reference data remain a critical bottleneck for crop-type mapping at any spatial and temporal scale. Operational systems such as WorldCereal aggregate labels from heterogeneous sources such as parcel registers, national databases, field surveys, and map-derived products, each with their own biases, coverage gaps and unknown label noise. Simple global rules are inadequate, since crop phenology and observation conditions vary strongly across regions and seasons. In this study, we focus on a single, operationally relevant question: whether embeddings produced through geospatial foundation models are a viable basis for cleaning the reference data. We propose a practical, locality-aware, embedding-based anomaly (EBA) detection framework that operates on the embeddings of a pretrained Earth-observation encoder. We score each labelled sample against other samples of the same crop in the same area using a pretrained embedding, flag the ones that stand out, and test whether removing or down-weighting them before training yields a better model. We establish that the flagged points are genuinely mislabelled or misplaced in two independent ways: against synthetic ground truth, the detector concentrates injected label errors 2.5-5x above chance in its flagged set (detection AUROC up to 0.84); and on real data, a model-independent test shows that removing or confidence-weighting the flagged held-out points raises measured accuracy in trained models, for both crop type and land cover. Acting on the flags then improves the WorldCereal crop-type model across five macro-regions, evaluated on a fixed held-out split under three views. We find conservative cleaning helps while over-cleaning hurts. The EBA detector approach is designed to be reproducible and extensible, and can serve as a template for cleaning large, noisy Earth observation reference datasets beyond crop mapping.

遥感数据清洗嵌入检测作物分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。