YOLOv11改进版实时检测混凝土裂缝,小目标识别更准。
A Real-time Concrete Crack Detection and Segmentation Model Based on YOLOv11
- 引入动态核仓库与三重注意力机制,增强特征提取能力。
- 在复杂背景下实现91.3%精度、76.6%召回率、86.4% mAP@50。
- 适合工程巡检场景,对数据少和噪声有强鲁棒性。
长三角地区交通基础设施加速老化,亟需高效混凝土裂缝检测以保障结构安全与经济发展。针对人工巡检效率低及现有深度学习模型在复杂背景中对小目标裂缝检测性能不佳的问题,本文提出基于YOLOv11n架构的多任务裂缝检测与分割模型YOLOv11-KW-TA-FP。该模型采用三阶段优化框架:(1)在主干网络中嵌入动态核仓库卷积(KWConv),通过动态核共享机制提升特征表示;(2)在特征金字塔中引入三重注意力机制(TA),强化通道-空间交互建模;(3)设计FP-IoU损失函数,实现自适应边界框回归惩罚。实验表明,改进模型相较基线显著提升性能,达到91.3%精度、76.6%召回率、86.4% mAP@50。消融实验证明各模块协同有效。鲁棒性测试显示,在数据稀缺与噪声干扰下仍保持稳定表现。本研究为自动化基础设施巡检提供高效视觉解决方案,具备显著工程应用价值。
原文摘要 · Abstract (English)
Accelerated aging of transportation infrastructure in the rapidly developing Yangtze River Delta region necessitates efficient concrete crack detection, as crack deterioration critically compromises structural integrity and regional economic growth. To overcome the limitations of inefficient manual inspection and the suboptimal performance of existing deep learning models, particularly for small-target crack detection within complex backgrounds, this paper proposes YOLOv11-KW-TA-FP, a multi-task concrete crack detection and segmentation model based on the YOLOv11n architecture. The proposed model integrates a three-stage optimization framework: (1) Embedding dynamic KernelWarehouse convolution (KWConv) within the backbone network to enhance feature representation through a dynamic kernel sharing mechanism; (2) Incorporating a triple attention mechanism (TA) into the feature pyramid to strengthen channel-spatial interaction modeling; and (3) Designing an FP-IoU loss function to facilitate adaptive bounding box regression penalization. Experimental validation demonstrates that the enhanced model achieves significant performance improvements over the baseline, attaining 91.3% precision, 76.6% recall, and 86.4% mAP@50. Ablation studies confirm the synergistic efficacy of the proposed modules. Furthermore, robustness tests indicate stable performance under conditions of data scarcity and noise interference. This research delivers an efficient computer vision solution for automated infrastructure inspection, exhibiting substantial practical engineering value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。