首个多类别水稻分割数据集,覆盖全球五大产稻国。
Global Rice Multi-Class Segmentation Dataset (RiceSEG): A Comprehensive and Diverse High-Resolution RGB-Annotated Images for the Development and Benchmarking of Rice Segmentation Algorithms
- 构建包含6000多种水稻基因型的高分辨率图像数据集
- 涵盖所有生长阶段,标注六类目标:绿叶、枯叶、穗、杂草等
- 适合开发精准农业中的水稻表型分析算法
开发基于计算机视觉的水稻表型技术对精准田间管理和加速育种至关重要。在表型任务中,区分图像成分是器官尺度植物生长发育表征的关键前提,有助于深入理解生态生理过程。然而,由于水稻器官结构精细且冠层光照复杂,该任务仍具挑战性,亟需高质量训练数据。现有数据集稀缺,既因缺乏大规模代表性水稻田图像,也因标注耗时。为此,我们建立了首个全面的多类别水稻语义分割数据集——RiceSEG。从中国、日本、印度、菲律宾和坦桑尼亚五个主要产稻国采集近5万张高分辨率地面图像,涵盖超过6000个基因型及所有生长阶段。从中精选3078张代表性样本,标注六个类别(背景、绿色植被、衰老植被、穗、杂草、鸭舌草)构成RiceSEG数据集。其中中国子数据集覆盖从东北到华南的主要基因型与种植环境。采用先进卷积神经网络与基于Transformer的分割模型作为基线。尽管在背景与绿色植被分割上表现尚可,但在生殖期因冠层结构复杂、多类别共存而面临困难。这些结果凸显了本数据集对研发专用水稻及其他作物分割模型的重要性。
原文摘要 · Abstract (English)
Developing computer vision-based rice phenotyping techniques is crucial for precision field management and accelerating breeding, thereby continuously advancing rice production. Among phenotyping tasks, distinguishing image components is a key prerequisite for characterizing plant growth and development at the organ scale, enabling deeper insights into eco-physiological processes. However, due to the fine structure of rice organs and complex illumination within the canopy, this task remains highly challenging, underscoring the need for a high-quality training dataset. Such datasets are scarce, both due to a lack of large, representative collections of rice field images and the time-intensive nature of annotation. To address this gap, we established the first comprehensive multi-class rice semantic segmentation dataset, RiceSEG. We gathered nearly 50,000 high-resolution, ground-based images from five major rice-growing countries (China, Japan, India, the Philippines, and Tanzania), encompassing over 6,000 genotypes across all growth stages. From these original images, 3,078 representative samples were selected and annotated with six classes (background, green vegetation, senescent vegetation, panicle, weeds, and duckweed) to form the RiceSEG dataset. Notably, the sub-dataset from China spans all major genotypes and rice-growing environments from the northeast to the south. Both state-of-the-art convolutional neural networks and transformer-based semantic segmentation models were used as baselines. While these models perform reasonably well in segmenting background and green vegetation, they face difficulties during the reproductive stage, when canopy structures are more complex and multiple classes are involved. These findings highlight the importance of our dataset for developing specialized segmentation models for rice and other crops.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。