arXiv:2412.14211cs.CVcs.LG2024-12被引 7

改进YOLOv8在相机陷阱数据上的泛化能力,提升野外识别准确率。

Improving Generalization Performance of YOLOv8 for Camera Trap Object Detection

  • 引入全局注意力机制与改进的多尺度融合,增强特征提取能力。
  • 使用WIoUv3损失函数,显著提升边界框回归精度。
  • 在新环境数据上表现更稳健,适合野外生态监测场景。

相机陷阱已成为野生动物保护的重要工具,可非侵入式地监测自然栖息地中的动物。利用目标检测算法自动识别相机陷阱图像中的物种对研究与保护至关重要。然而,模型在未见过的数据集上泛化能力差的问题普遍存在。本文针对基准YOLOv8模型在真实环境中的泛化缺陷,提出三项改进:引入全局注意力机制(GAM)模块、优化多尺度特征融合策略,并采用Wise IoU v3(WIoUv3)作为边界框回归损失函数。通过全面评估与消融实验验证,改进后的模型能有效抑制背景噪声,聚焦物体关键特征,在新环境数据上展现出更强的鲁棒性。该方法不仅缓解了相机陷阱数据集的固有挑战,也为实际生态保护应用提供了更可靠的解决方案,助力野生动物种群与栖息地的高效管理。

原文摘要 · Abstract (English)

Camera traps have become integral tools in wildlife conservation, providing non-intrusive means to monitor and study wildlife in their natural habitats. The utilization of object detection algorithms to automate species identification from Camera Trap images is of huge importance for research and conservation purposes. However, the generalization issue, where the trained model is unable to apply its learnings to a never-before-seen dataset, is prevalent. This thesis explores the enhancements made to the YOLOv8 object detection algorithm to address the problem of generalization. The study delves into the limitations of the baseline YOLOv8 model, emphasizing its struggles with generalization in real-world environments. To overcome these limitations, enhancements are proposed, including the incorporation of a Global Attention Mechanism (GAM) module, modified multi-scale feature fusion, and Wise Intersection over Union (WIoUv3) as a bounding box regression loss function. A thorough evaluation and ablation experiments reveal the improved model's ability to suppress the background noise, focus on object properties, and exhibit robust generalization in novel environments. The proposed enhancements not only address the challenges inherent in camera trap datasets but also pave the way for broader applicability in real-world conservation scenarios, ultimately aiding in the effective management of wildlife populations and habitats.

目标检测相机陷阱YOLOv8泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。