arXiv:2509.20856cs.CV2025-09被引 105

用网络噪声数据训练植物识别模型,效果竟超过专家标注数据。

Plant identification based on noisy web data: the amazing performance of deep learning (LifeCLEF 2017)

  • 用网页收集的含错标签数据训练模型,替代人工精标数据。
  • 在100万张图上训练,识别准确率达85.3%(测试集来自Pl@ntNet)。
  • 适合做大规模植物识别系统的研究者与生态数据开发者。

2017年LifeCLEF植物识别挑战赛是迈向可覆盖欧北美大陆1万种植物的自动化识别系统的重要里程碑,共包含110万张图像。尽管有像EOL这样的国际项目整合了权威机构的植物图像,但多数物种仍缺乏图像或图文质量差。而网络上存在大量由植物爱好者、博客、图片站和电商网站提供的植物图像,虽含大量错误标签,但数量庞大。本挑战赛旨在评估这类大规模噪声训练数据是否能媲美小规模但专家校验的可信数据。为公平比较,测试集来自Pl@ntNet移动端应用全球采集的数百万张查询图像。本文详细介绍了挑战资源与评估方式,总结了参赛团队的方法与系统,并分析了主要结果。

原文摘要 · Abstract (English)

The 2017-th edition of the LifeCLEF plant identification challenge is an important milestone towards automated plant identification systems working at the scale of continental floras with 10.000 plant species living mainly in Europe and North America illustrated by a total of 1.1M images. Nowadays, such ambitious systems are enabled thanks to the conjunction of the dazzling recent progress in image classification with deep learning and several outstanding international initiatives, such as the Encyclopedia of Life (EOL), aggregating the visual knowledge on plant species coming from the main national botany institutes. However, despite all these efforts the majority of the plant species still remain without pictures or are poorly illustrated. Outside the institutional channels, a much larger number of plant pictures are available and spread on the web through botanist blogs, plant lovers web-pages, image hosting websites and on-line plant retailers. The LifeCLEF 2017 plant challenge presented in this paper aimed at evaluating to what extent a large noisy training dataset collected through the web and containing a lot of labelling errors can compete with a smaller but trusted training dataset checked by experts. To fairly compare both training strategies, the test dataset was created from a third data source, i.e. the Pl@ntNet mobile application that collects millions of plant image queries all over the world. This paper presents more precisely the resources and assessments of the challenge, summarizes the approaches and systems employed by the participating research groups, and provides an analysis of the main outcomes.

植物识别噪声数据深度学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。