YOLO-FDA提升工业表面缺陷检测精度与鲁棒性
YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection

- 引入细节导向融合模块和注意力特征融合机制
- 在多个数据集上达到更高检测准确率和多尺度稳定性
- 适合工业质检中细小、复杂纹理缺陷的识别任务
工业场景中的表面缺陷检测既关键又具挑战性,因缺陷类型多样、形状尺寸不规则、细节要求精细且材质纹理复杂。尽管基于AI的检测器性能有所提升,现有方法仍存在特征冗余、细节敏感度不足、多尺度下鲁棒性弱等问题。为此,本文提出YOLO-FDA,一种集成细粒度细节增强与注意力引导特征融合的新型YOLO框架。具体地,采用类似BiFPN的架构强化YOLOv5主干网络内的双向多层级特征聚合;为更好捕捉细微结构变化,引入第二低层的方向性异构卷积(DDFM),丰富空间细节,并将该层与低层特征融合以增强语义一致性。此外,提出两种新颖的注意力融合策略:注意力加权拼接(AC)与跨层注意力融合(CAF),以改善上下文表征并降低特征噪声。大量实验表明,YOLO-FDA在多个基准数据集上持续优于现有最先进方法,在各类缺陷及不同尺度下均表现出更高准确率与更强鲁棒性。
原文摘要 · Abstract (English)
Surface defect detection in industrial scenarios is both crucial and technically demanding due to the wide variability in defect types, irregular shapes and sizes, fine-grained requirements, and complex material textures. Although recent advances in AI-based detectors have improved performance, existing methods often suffer from redundant features, limited detail sensitivity, and weak robustness under multiscale conditions. To address these challenges, we propose YOLO-FDA, a novel YOLO-based detection framework that integrates fine-grained detail enhancement and attention-guided feature fusion. Specifically, we adopt a BiFPN-style architecture to strengthen bidirectional multilevel feature aggregation within the YOLOv5 backbone. To better capture fine structural changes, we introduce a Detail-directional Fusion Module (DDFM) that introduces a directional asymmetric convolution in the second-lowest layer to enrich spatial details and fuses the second-lowest layer with low-level features to enhance semantic consistency. Furthermore, we propose two novel attention-based fusion strategies, Attention-weighted Concatenation (AC) and Cross-layer Attention Fusion (CAF) to improve contextual representation and reduce feature noise. Extensive experiments on benchmark datasets demonstrate that YOLO-FDA consistently outperforms existing state-of-the-art methods in terms of both accuracy and robustness across diverse types of defects and scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。