arXiv:2503.20068cs.CV2025-03CVPR被引 4

构建470万张图像的农作物与杂草识别数据集,推动精准农业发展。

iNatAg: Multi-Class Classification Models Enabled by a Large-Scale Benchmark Dataset with 4.7M Images of 2,959 Crop and Weed Species

论文配图:iNatAg: Multi-Class Classification Models Enabled by a Large-Scale Benchmark Dataset with 4.7M Images of 2,959 Crop and Weed Species
图 1 · 摘自论文原文
  • 基于Swin Transformer构建多任务分类模型,融合地理信息与LoRA微调。
  • 在作物与杂草分类上达到92.38%准确率,领先当前最优水平。
  • 适合从事农业视觉识别、植物分类与智能耕作系统研发者使用。

准确识别作物与杂草种类对精准农业和可持续耕作至关重要。然而,由于物种间视觉相似性高、环境变化大以及农业专用图像数据稀缺,该任务仍具挑战。本文提出iNatAg,一个包含470万张图像、覆盖2,959种作物与杂草的大型图像数据集,具备从二元分类到具体物种级别的精确分类标注。数据源自iNaturalist,覆盖全球各大洲,真实反映自然图像与环境多样性。基于此数据集,我们训练了基于Swin Transformer的基准模型,并评估了引入地理空间数据与LoRA微调等改进方法的效果。最佳模型在所有分类任务中均达到当前最优性能,作物与杂草分类准确率达92.38%。数据规模使我们得以分析误分类模式,开启植物分类新分析路径。通过大规模物种覆盖、多任务标签与地理多样性,iNatAg为构建鲁棒、地理位置感知的农业分类系统提供新基础。数据集已通过AgML公开发布(https://github.com/Project-AgML/AgML),支持直接接入农业机器学习工作流。

原文摘要 · Abstract (English)

Accurate identification of crop and weed species is critical for precision agriculture and sustainable farming. However, it remains a challenging task due to a variety of factors -- a high degree of visual similarity among species, environmental variability, and a continued lack of large, agriculture-specific image data. We introduce iNatAg, a large-scale image dataset which contains over 4.7 million images of 2,959 distinct crop and weed species, with precise annotations along the taxonomic hierarchy from binary crop/weed labels to specific species labels. Curated from the broader iNaturalist database, iNatAg contains data from every continent and accurately reflects the variability of natural image captures and environments. Enabled by this data, we train benchmark models built upon the Swin Transformer architecture and evaluate the impact of various modifications such as the incorporation of geospatial data and LoRA finetuning. Our best models achieve state-of-the-art performance across all taxonomic classification tasks, achieving 92.38\% on crop and weed classification. Furthermore, the scale of our dataset enables us to explore incorrect misclassifications and unlock new analytic possiblities for plant species. By combining large-scale species coverage, multi-task labels, and geographic diversity, iNatAg provides a new foundation for building robust, geolocation-aware agricultural classification systems. We release the iNatAg dataset publicly through AgML (https://github.com/Project-AgML/AgML), enabling direct access and integration into agricultural machine learning workflows.

农业视觉植物识别多分类数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。