arXiv:2501.17983cs.CV2025-01被引 19

提升无人机图像小目标检测精度,兼顾定位与分类性能。

Efficient Feature Fusion for UAV Object Detection

  • 设计混合上采样与下采样模块,灵活调整多层特征图分辨率。
  • 在两个公开数据集上,相比基线模型提升2%平均精度(AP)。
  • 适合需要高精度小目标检测的无人机遥感应用。

无人机遥感图像中的目标检测面临图像质量不稳定、目标尺寸小、背景复杂及环境遮挡等挑战。小目标占据图像比例小,检测难度大。现有多尺度特征融合方法虽部分缓解此问题,但难以平衡分类与定位性能,主要因特征表示不足和网络信息流不平衡。本文提出一种专为无人机目标检测设计的新颖特征融合框架,同时提升定位精度与分类能力。该框架集成混合上采样与下采样模块,可将不同深度的特征图灵活调整至任意分辨率,促进跨层连接与多尺度融合,增强小目标表征。下采样模块强化细粒度特征,提升复杂条件下的空间定位能力;上采样模块聚合全局上下文信息,优化多尺度特征一致性,增强杂乱场景中的分类鲁棒性。在两个公开无人机数据集上的实验表明,该框架集成于YOLO-v10模型后,相比基准模型实现2%的平均精度(AP)提升,参数量保持不变。结果验证了该框架在精准高效无人机目标检测中的潜力。

原文摘要 · Abstract (English)

Object detection in unmanned aerial vehicle (UAV) remote sensing images poses significant challenges due to unstable image quality, small object sizes, complex backgrounds, and environmental occlusions. Small objects, in particular, occupy small portions of images, making their accurate detection highly difficult. Existing multi-scale feature fusion methods address these challenges to some extent by aggregating features across different resolutions. However, they often fail to effectively balance the classification and localization performance for small objects, primarily due to insufficient feature representation and imbalanced network information flow. In this paper, we propose a novel feature fusion framework specifically designed for UAV object detection tasks to enhance both localization accuracy and classification performance. The proposed framework integrates hybrid upsampling and downsampling modules, enabling feature maps from different network depths to be flexibly adjusted to arbitrary resolutions. This design facilitates cross-layer connections and multi-scale feature fusion, ensuring improved representation of small objects. Our approach leverages hybrid downsampling to enhance fine-grained feature representation, improving spatial localization of small targets, even under complex conditions. Simultaneously, the upsampling module aggregates global contextual information, optimizing feature consistency across scales and enhancing classification robustness in cluttered scenes. Experimental results on two public UAV datasets demonstrate the effectiveness of the proposed framework. Integrated into the YOLO-v10 model, our method achieves a 2% improvement in average precision (AP) compared to the baseline YOLO-v10 model, while maintaining the same number of parameters. These results highlight the potential of our framework for accurate and efficient UAV object detection.

无人机检测小目标特征融合YOLO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。