用遥感数据预测欧洲植物分布,挑战多标签分类难题
Overview of GeoLifeCLEF 2023: Species Composition Prediction with High Spatial Resolution at Continental Scale Using Remote Sensing
- 结合遥感与环境数据,用深度学习建模物种分布
- 在2.2万个小样方上实现大陆尺度物种组成预测
- 提出单标签与多标签数据联合训练新策略
理解物种的时空分布是生态学与保护研究的核心。通过将物种观测与地理及环境变量结合,可建立环境与可能存在的物种之间的关系模型。为推动该领域在深度学习与遥感数据方面的进展,我们组织了名为GeoLifeCLEF 2023的开放机器学习竞赛。训练数据包含500万条植物物种观测记录(每样本单正类标签),覆盖欧洲大部分植被区,以及高分辨率遥感影像、土地覆盖、高程数据,辅以气候、土壤和人类足迹等粗分辨率变量。本任务为多标签分类,评估模型基于标准化调查,在2.2万个小型样方上预测物种组成的能力。本文综述竞赛情况,总结参赛团队的方法并分析主要结果。特别指出:采用单正类标签训练的方法在多标签评估中存在偏差,并提出一种有效融合单标签与多标签数据的新型训练策略。
原文摘要 · Abstract (English)
Understanding the spatio-temporal distribution of species is a cornerstone of ecology and conservation. By pairing species observations with geographic and environmental predictors, researchers can model the relationship between an environment and the species which may be found there. To advance the state-of-the-art in this area with deep learning models and remote sensing data, we organized an open machine learning challenge called GeoLifeCLEF 2023. The training dataset comprised 5 million plant species observations (single positive label per sample) distributed across Europe and covering most of its flora, high-resolution rasters: remote sensing imagery, land cover, elevation, in addition to coarse-resolution data: climate, soil and human footprint variables. In this multi-label classification task, we evaluated models ability to predict the species composition in 22 thousand small plots based on standardized surveys. This paper presents an overview of the competition, synthesizes the approaches used by the participating teams, and analyzes the main results. In particular, we highlight the biases faced by the methods fitted to single positive labels when it comes to the multi-label evaluation, and the new and effective learning strategy combining single and multi-label data in training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。