用标本馆资料提升热带植物自动识别,挑战跨域分类难题。
Overview of LifeCLEF Plant Identification task 2020
- 用数百万张标本图+少量野外照片训练模型,实现跨域识别
- 测试集全为野外照片,验证在数据匮乏地区的效果
- 适合关注生物多样性保护与跨域图像识别的研究者
由于深度学习进展和野外照片数据增多,植物自动识别能力显著提升。然而这些数据主要集中于北美洲和西欧,仅覆盖数万种植物,对生物多样性最丰富的热带地区覆盖不足。相反,数百年的植物标本被系统存于标本馆,尤其在热带地区,近年数字化使数百万张标本图在线可得。2020年LifeCLEF植物识别挑战赛(PlantCLEF 2020)旨在评估利用标本馆数据提升数据匮乏地区植物识别的效果。该任务基于约1,000种植物的数据集,主要聚焦南美洲圭亚那高地这一全球植物多样性最高的区域之一。挑战赛采用跨域分类形式:训练集包含数十万张标本图和数千张野外照片,用于学习两个域间的映射;测试集则完全由野外照片构成。本文介绍了所用资源与评估方法,总结了参赛团队的技术方案,并分析了主要结果。
原文摘要 · Abstract (English)
Automated identification of plants has improved considerably thanks to the recent progress in deep learning and the availability of training data with more and more photos in the field. However, this profusion of data only concerns a few tens of thousands of species, mostly located in North America and Western Europe, much less in the richest regions in terms of biodiversity such as tropical countries. On the other hand, for several centuries, botanists have collected, catalogued and systematically stored plant specimens in herbaria, particularly in tropical regions, and the recent efforts by the biodiversity informatics community made it possible to put millions of digitized sheets online. The LifeCLEF 2020 Plant Identification challenge (or "PlantCLEF 2020") was designed to evaluate to what extent automated identification on the flora of data deficient regions can be improved by the use of herbarium collections. It is based on a dataset of about 1,000 species mainly focused on the South America's Guiana Shield, an area known to have one of the greatest diversity of plants in the world. The challenge was evaluated as a cross-domain classification task where the training set consist of several hundred thousand herbarium sheets and few thousand of photos to enable learning a mapping between the two domains. The test set was exclusively composed of photos in the field. This paper presents the resources and assessments of the conducted evaluation, summarizes the approaches and systems employed by the participating research groups, and provides an analysis of the main outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。