arXiv:2410.07955cs.CV2024-10被引 10

用迭代标注与轻量模型提升无人机香蕉园图像分割效率

Iterative Optimization Annotation Pipeline and ALSS-YOLO-Seg for Efficient Banana Plantation Segmentation in UAV Imagery

  • 用SAM2零样本能力迭代优化标注,大幅降低人工标注成本
  • 新模型ALSS-YOLO-Seg在复杂背景下实现高精度作物分割
  • 适合农业遥感、边缘设备部署的高效分割任务

无人机拍摄的香蕉园图像精确分割对产量预估和植株健康评估至关重要。通过识别与分类种植区域可计算作物面积,这对精准预测不可或缺。然而,香蕉园图像分割需大量标注数据,人工标注耗时费力,限制了大规模数据集的发展。此外,目标尺寸变化大、地面背景复杂、计算资源有限及作物类别准确识别等问题使分割难度加剧。为此,我们提出综合解决方案:首先设计基于SAM2零样本能力的迭代优化标注流程,显著降低标注成本与时间;其次开发面向无人机影像的高效轻量级分割模型ALSS-YOLO-Seg。该模型主干引入自适应轻量通道分裂与混洗(ALSS)模块,增强通道间信息交换并优化特征提取,提升作物识别精度;同时结合多尺度通道注意力(MSCA)模块,融合多尺度特征提取与通道注意力机制,有效应对目标尺寸变化与复杂背景挑战。

原文摘要 · Abstract (English)

Precise segmentation of Unmanned Aerial Vehicle (UAV)-captured images plays a vital role in tasks such as crop yield estimation and plant health assessment in banana plantations. By identifying and classifying planted areas, crop area can be calculated, which is indispensable for accurate yield predictions. However, segmenting banana plantation scenes requires a substantial amount of annotated data, and manual labeling of these images is both time-consuming and labor-intensive, limiting the development of large-scale datasets. Furthermore, challenges such as changing target sizes, complex ground backgrounds, limited computational resources, and correct identification of crop categories make segmentation even more difficult. To address these issues, we proposed a comprehensive solution. Firstly, we designed an iterative optimization annotation pipeline leveraging SAM2's zero-shot capabilities to generate high-quality segmentation annotations, thereby reducing the cost and time associated with data annotation significantly. Secondly, we developed ALSS-YOLO-Seg, an efficient lightweight segmentation model optimized for UAV imagery. The model's backbone includes an Adaptive Lightweight Channel Splitting and Shuffling (ALSS) module to improve information exchange between channels and optimize feature extraction, aiding accurate crop identification. Additionally, a Multi-Scale Channel Attention (MSCA) module combines multi-scale feature extraction with channel attention to tackle challenges of varying target sizes and complex ground backgrounds.

图像分割无人机影像农业遥感轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。