arXiv:2601.16617cs.CVcs.AI2026-01

通过挖掘边界与位置信息提升无人机小目标检测精度

Boundary and Position Information Mining for Aerial Small Object Detection

  • 引入位置与边界引导模块,融合多尺度特征增强定位能力
  • 在VisDrone2021等数据集上显著优于Yolov5-P2基线
  • 适合无人机、遥感等小目标密集场景的检测任务

无人机应用在航拍与目标识别中日益普及,但因尺度不平衡和边缘模糊,小目标检测面临挑战。为此提出边界与位置信息挖掘(BPIM)框架,包含位置信息引导(PIG)、边界信息引导(BIG)、跨尺度融合(CSF)、三重特征融合(TFF)及自适应权重融合(AWF)模块。该框架利用注意力机制与跨尺度特征融合策略,整合图像中的边界、位置与尺度信息。实验表明,在VisDrone2021、DOTA1.0和WiderPerson数据集上,BPIM性能优于基线Yolov5-P2,且在计算开销相当的情况下达到当前先进水平。

原文摘要 · Abstract (English)

Unmanned Aerial Vehicle (UAV) applications have become increasingly prevalent in aerial photography and object recognition. However, there are major challenges to accurately capturing small targets in object detection due to the imbalanced scale and the blurred edges. To address these issues, boundary and position information mining (BPIM) framework is proposed for capturing object edge and location cues. The proposed BPIM includes position information guidance (PIG) module for obtaining location information, boundary information guidance (BIG) module for extracting object edge, cross scale fusion (CSF) module for gradually assembling the shallow layer image feature, three feature fusion (TFF) module for progressively combining position and boundary information, and adaptive weight fusion (AWF) module for flexibly merging the deep layer semantic feature. Therefore, BPIM can integrate boundary, position, and scale information in image for small object detection using attention mechanisms and cross-scale feature fusion strategies. Furthermore, BPIM not only improves the discrimination of the contextual feature by adaptive weight fusion with boundary, but also enhances small object perceptions by cross-scale position fusion. On the VisDrone2021, DOTA1.0, and WiderPerson datasets, experimental results show the better performances of BPIM compared to the baseline Yolov5-P2, and obtains the promising performance in the state-of-the-art methods with comparable computation load.

小目标检测无人机视觉边界感知多尺度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。