构建5000张葡萄园图像数据集,实现葡萄簇闭合度自动精准评估
ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation

- 用点定位+提示分割+变压器架构,实现高精度葡萄粒定位与簇分割
- 覆盖13个品种,标注超64万颗葡萄粒,支持复杂场景下的闭合度估算
- 适合农业智能监测、作物表型分析研究者使用
葡萄簇闭合度指果粒间缝隙逐渐填充的程度,是葡萄园管理中的关键性状,影响病害风险。传统人工目测方法费时、主观且缺乏时间分辨率。现有数据集极少支持细粒度葡萄粒级分析,制约深度学习模型发展。本文提出大规模田间葡萄视觉数据集ViViD-5K,包含5000张图像,涵盖超过648,000个葡萄粒中心点及簇分割掩码,覆盖13种葡萄品种。基于此,我们提出GrapeSAM,一种两阶段视觉流程:先通过点定位进行葡萄粒检测,再结合Segment Anything的提示分割与变压器结构完成簇级分割。该方法可在少量监督下实现田间自动闭合度估计。定量结果显示在多种条件下均具备优异分割与计数精度,可视化验证了对域内及域外样本的鲁棒性。本工作为人工紧凑度评分提供了可扩展、客观的替代方案,支持高通量葡萄表型分析并提升空间细节表现。
原文摘要 · Abstract (English)
Cluster closure, defined as the progressive filling of gaps between the berries in a grape bunch, is a key trait in vineyard management, impacting disease risk. However, traditional visual scoring methods are labor-intensive, subjective, and lack temporal resolution. Existing datasets rarely support fine-grained berry-level analysis, limiting the development of robust deep learning models. In this work, we present ViViD-5k, a large-scale in-field Vineyard Vision Dataset containing 5,000 images with dense annotations, including over 648,000 berry centroids and cluster segmentation masks spanning 13 grape varieties. Building on this dataset, we introduce GrapeSAM, a two-stage visual pipeline that combines point-based berry localization with prompt-based segmentation using Segment Anything, followed by transformer-based cluster segmentation. The pipeline enables automated, in-field estimation of cluster closure with minimal supervision. Quantitative results demonstrate strong segmentation and counting accuracy across diverse conditions, while visualizations confirm robustness on both in-domain and out-of-domain samples. This work provides a scalable and objective alternative to manual compactness scoring and supports high-throughput grape phenotyping with enhanced spatial detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。