提出作物对齐遮挡法,提升农田导航模型鲁棒性
CA-Cut: Crop-Aligned Cutout for Data Augmentation to Learn More Robust Under-Canopy Navigation
- 用作物行位置引导遮挡区域,增强模型上下文理解能力
- 在玉米田数据集上使预测误差降低36.9%
- 适合需要强泛化能力的农业视觉导航研究
当前先进的视觉下冠导航方法依赖深度学习感知模型区分可通行区域与作物行。尽管性能优异,但需大量训练数据以确保真实田间部署的可靠性。然而,数据采集成本高昂,需大量人力进行实地采样与标注。为此,训练中常采用色彩抖动、高斯模糊、水平翻转等数据增强技术以扩充数据多样性并提升模型鲁棒性。本文假设仅使用这些常规增强手段在复杂下冠环境(频繁遮挡、杂物、作物间距不均)中效果有限。为此,提出新型增强方法Crop-Aligned Cutout(CA-Cut),在输入图像中随机遮挡位于作物行侧边的空间区域,促使模型在细粒度信息被遮蔽时仍能捕捉高层上下文特征。在公开玉米田数据集上的大量实验表明,基于遮挡的增强能有效模拟遮挡,显著提升语义关键点预测的鲁棒性。尤其发现,将遮挡分布偏向作物行可显著提升预测准确率与跨环境泛化能力,最高实现36.9%的预测误差下降。此外,通过消融实验确定了遮挡数量、大小及空间分布的最优配置以最大化整体性能。
原文摘要 · Abstract (English)
State-of-the-art visual under-canopy navigation methods are designed with deep learning-based perception models to distinguish traversable space from crop rows. While these models have demonstrated successful performance, they require large amounts of training data to ensure reliability in real-world field deployment. However, data collection is costly, demanding significant human resources for in-field sampling and annotation. To address this challenge, various data augmentation techniques are commonly employed during model training, such as color jittering, Gaussian blur, and horizontal flip, to diversify training data and enhance model robustness. In this paper, we hypothesize that utilizing only these augmentation techniques may lead to suboptimal performance, particularly in complex under-canopy environments with frequent occlusions, debris, and non-uniform spacing of crops. Instead, we propose a novel augmentation method, so-called Crop-Aligned Cutout (CA-Cut) which masks random regions out in input images that are spatially distributed around crop rows on the sides to encourage trained models to capture high-level contextual features even when fine-grained information is obstructed. Our extensive experiments with a public cornfield dataset demonstrate that masking-based augmentations are effective for simulating occlusions and significantly improving robustness in semantic keypoint predictions for visual navigation. In particular, we show that biasing the mask distribution toward crop rows in CA-Cut is critical for enhancing both prediction accuracy and generalizability across diverse environments achieving up to a 36.9% reduction in prediction error. In addition, we conduct ablation studies to determine the number of masks, the size of each mask, and the spatial distribution of masks to maximize overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。