提出可自适应去除背景干扰的实时道路损伤检测模型
Real-Time Dynamic Scale-Aware Fusion Detection Network: Take Road Damage Detection as an example
- 设计多尺度自适应融合模块,提升对不规则损伤的感知能力
- 在UAV-PDD2023上达到54.2% mAP50,比YOLOv10-m高11.1%
- 参数量仅1.8M,推理速度更快,适合部署在无人机端
基于无人机的道路损伤检测(RDD)对城市日常维护与安全至关重要,能显著降低人力成本。然而,现有研究仍面临诸多挑战:损伤尺寸和方向不规则、背景遮挡以及与背景难以区分等问题,严重影响无人机在日常巡检中的检测能力。为解决上述问题并提升无人机实时道路损伤检测性能,本文提出三种对应模块:可灵活适应形状与背景的特征提取模块;融合多尺度感知并自适应形状与背景的融合模块;高效下采样模块。基于这些模块,构建了具备自动去除背景干扰能力的多尺度自适应道路损伤检测模型——动态尺度感知融合检测模型(RT-DSAFDet)。在UAV-PDD2023公开数据集上的实验结果表明,该模型在保持轻量化的同时,实现了54.2%的mAP50,较YOLOv10-m提升11.1%,参数量降至1.8M,FLOPs降至4.6G,分别减少88%和93%。此外,在大规模通用目标检测数据集MS COCO2017上,其mAP50-95与YOLOv9-t相当,但mAP50高出0.5%,参数量减少10%,计算量降低40%。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicle (UAV)-based Road Damage Detection (RDD) is important for daily maintenance and safety in cities, especially in terms of significantly reducing labor costs. However, current UAV-based RDD research is still faces many challenges. For example, the damage with irregular size and direction, the masking of damage by the background, and the difficulty of distinguishing damage from the background significantly affect the ability of UAV to detect road damage in daily inspection. To solve these problems and improve the performance of UAV in real-time road damage detection, we design and propose three corresponding modules: a feature extraction module that flexibly adapts to shape and background; a module that fuses multiscale perception and adapts to shape and background ; an efficient downsampling module. Based on these modules, we designed a multi-scale, adaptive road damage detection model with the ability to automatically remove background interference, called Dynamic Scale-Aware Fusion Detection Model (RT-DSAFDet). Experimental results on the UAV-PDD2023 public dataset show that our model RT-DSAFDet achieves a mAP50 of 54.2%, which is 11.1% higher than that of YOLOv10-m, an efficient variant of the latest real-time object detection model YOLOv10, while the amount of parameters is reduced to 1.8M and FLOPs to 4.6G, with a decreased by 88% and 93%, respectively. Furthermore, on the large generalized object detection public dataset MS COCO2017 also shows the superiority of our model with mAP50-95 is the same as YOLOv9-t, but with 0.5% higher mAP50, 10% less parameters volume, and 40% less FLOPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。