针对弹簧表面缺陷检测难题,提出多尺度可变形注意力框架,兼顾精度与实时性。
A Deformable Attention-Based Detection Transformer with Cross-Scale Feature Fusion for Industrial Coil Spring Inspection
- 采用结构重参数化设计,训练时多分支增强特征提取,推理时高效单路加速。
- 引入可变形注意力机制,动态聚焦缺陷区域,适应复杂形态和尺度变化。
- 融合跨尺度特征并结合GSConv与VoVGSCSP模块,提升小目标检测能力,适合工业质检场景。
机车弹簧自动化视觉检测因表面缺陷形态多样、尺度差异大及工业背景复杂而面临挑战。本文提出MSD-DETR(多尺度可变形检测Transformer)框架,通过三项创新应对:(1) 结构重参数化策略,将训练阶段的多分支拓扑与推理阶段的效率解耦,提升特征提取能力同时保持实时性能;(2) 可变形注意力机制,实现内容自适应的空间采样,动态聚焦缺陷相关区域,不受形态不规则影响;(3) 融合GSConv模块与VoVGSCSP块的跨尺度特征融合架构,有效聚合多分辨率信息。在真实机车弹簧数据集上的实验证明,MSD-DETR在98 FPS下达到92.4% [email protected],优于YOLOv8(+3.1% mAP)和基线RT-DETR(+2.8% mAP),且推理速度相当,树立了工业弹簧质量检测新基准。
原文摘要 · Abstract (English)
Automated visual inspection of locomotive coil springs presents significant challenges due to the morphological diversity of surface defects, substantial scale variations, and complex industrial backgrounds. This paper proposes MSD-DETR (Multi-Scale Deformable Detection Transformer), a novel detection framework that addresses these challenges through three key innovations: (1) a structural re-parameterization strategy that decouples training-time multi-branch topology from inference-time efficiency, enhancing feature extraction while maintaining real-time performance; (2) a deformable attention mechanism that enables content-adaptive spatial sampling, allowing dynamic focus on defect-relevant regions regardless of morphological irregularity; and (3) a cross-scale feature fusion architecture incorporating GSConv modules and VoVGSCSP blocks for effective multi-resolution information aggregation. Comprehensive experiments on a real-world locomotive coil spring dataset demonstrate that MSD-DETR achieves 92.4\% [email protected] at 98 FPS, outperforming state-of-the-art detectors including YOLOv8 (+3.1\% mAP) and the baseline RT-DETR (+2.8\% mAP) while maintaining comparable inference speed, establishing a new benchmark for industrial coil spring quality inspection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。