arXiv:2608.25465cs.CVcs.AI2026-08

用偏振成像和Transformer模型提升焊接缝自动分割的工业质检鲁棒性

Automatic weld seam segmentation for industrial quality control: a comparison of RGB and polarimetric imaging with CNN and transformer architectures

论文配图:Automatic weld seam segmentation for industrial quality control: a comparison of RGB and polarimetric imaging with CNN and transformer architectures
图 1 · 摘自论文原文
  • 对比RGB与偏振图像,结合CNN与Transformer架构进行焊缝分割
  • 偏振图像在非受控环境下实现0.93的mAP50,优于传统RGB的0.22-0.48
  • Transformer对视角变化更鲁棒,适合复杂工业现场应用

工业焊接件的视觉检测仍是自动化程度最低的环节,依赖人工经验导致结果差异大。本研究评估了从RGB与偏振图像中自动分割焊缝的可行性,对比了受控实验室条件与真实非受控环境下的表现。在统一的阈值无关协议下,比较了卷积神经网络(CNN)与基于Transformer的架构。受控条件下,CNN的平均掩码mAP50可达0.87,但在非受控环境下降至0.22–0.48,表明采集条件是系统性能的关键因素。采用保持几何一致性的增强策略,偏振成像实现了高达0.93的平均掩码mAP50,可在非受控环境中达到与最佳受控RGB相当的精度,无需控制采集条件。在视角变化测试中,尽管在分布内两者性能相近,但当测试时视角发生偏移时,所有CNN模型性能崩溃,而Transformer模型(尤其是RF-DETR)仍保持高精度。该差距在三个随机种子和分辨率匹配对照下均成立,说明是架构而非训练分辨率所致。在CNN中,容量提升在考虑种子方差后无法带来稳定分布内增益:固定视角下小型CNN已足够,而可变视角则需Transformer。

原文摘要 · Abstract (English)

Visual inspection of welded assemblies remains one of the least automated stages in many industrial production processes, still depending largely on the experience of human operators and thus subject to inter-operator variability; the manufacturing of special-purpose machinery cabins, the setting of this study, is one representative case. This work evaluates the feasibility of automatic weld seam segmentation from RGB and polarimetric imagery, comparing controlled laboratory acquisitions with images captured under real, uncontrolled conditions. Convolutional neural network (CNN) architectures and transformer-based architectures are benchmarked under a unified, threshold-independent protocol, training each CNN with three random seeds to separate genuine effects from seed noise. In controlled RGB conditions, CNN models reach a mean mask mAP50 of up to 0.87, but drop to 0.22-0.48 under uncontrolled acquisition, showing that the acquisition setup is a first-order component of the inspection system. Polarimetric imaging with alignment-preserving geometric augmentation localizes previously unseen welds with a mean mask mAP50 up to 0.93: on par with, rather than ahead of, the best controlled-RGB result, but reaching that accuracy on uncontrolled RGB without requiring acquisition control. The clearest architectural finding concerns viewpoint robustness. In-distribution, transformers and CNNs are broadly comparable; but under a test-time viewpoint shift, the transformer models, and RF-DETR in particular, retain high accuracy while every CNN collapses. The gap holds across three seeds and a resolution-matched control, pointing to architecture rather than training resolution. Within the CNN family, capacity brings no reliable in-distribution gain once seed variance is accounted for: small CNNs suffice for fixed viewpoints, transformers for variable ones.

焊缝分割偏振成像Transformer工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。