arXiv:2601.03637cs.CV2026-01被引 1

用可控生成技术解决裂缝分割数据少、跨域差的问题。

CrackSegFlow: Controllable Flow Matching Synthesis for Generalizable Crack Segmentation with a 50K Image-Mask Benchmark

  • 通过流匹配生成带像素对齐的裂缝图像,保持细结构连续性。
  • 合成数据使同域性能提升5.37 mIoU,跨域迁移提升13.12 mIoU。
  • 适合需要高泛化能力的工业缺陷检测研究者使用。

缺陷分割是基础设施在建设与运营阶段基于计算机视觉检测的核心任务。然而,由于像素级标注稀缺以及环境间的域偏移,实际部署受限。本文提出CrackSegFlow,一种可控制的流匹配合成方法,能从掩码生成像素对齐的裂缝图像。渲染器结合拓扑保持掩码注入与边缘门控,确保细长结构连续性。类别条件流匹配采样多样拓扑掩码,再生成对应的真实图像。此外,将裂缝注入无裂纹背景以增加干扰因素,降低误检率。在五个数据集上,使用CNN-Transformer主干网络,添加合成样本使同域性能提升5.37 mIoU和5.13 F1;目标引导的跨域合成(基于目标掩码统计)进一步带来13.12 mIoU和14.82 F1提升。我们还发布了包含50,000张图像-掩码对的基准数据集CSF-50K。

原文摘要 · Abstract (English)

Defect segmentation is central to computer vision based inspection of infrastructure assets during both construction and operation. However, deployment remains limited due to scarce pixel-level labels and domain shift across environments. We introduce CrackSegFlow, a controllable Flow Matching synthesis method that renders synthetic images of cracks from masks with pixel-level alignment. Our renderer combines topology-preserving mask injection with edge gating to maintain thin-structure continuity. Class-conditional FM samples masks for topology diversity, and CrackSegFlow renders aligned ground truth images from them. We further inject cracks onto crack-free backgrounds to diversify confounders and reduce false positives. Across five datasets and using a CNN-Transformer backbone, our results demonstrate that adding synthesized pairs improves in-domain performance by +5.37 mIoU and +5.13 F1, while target-guided cross-domain synthesis driven by target mask statistics adds +13.12 mIoU and +14.82 F1. We also release CSF-50K, a benchmark dataset comprising 50,000 image-mask pairs.

缺陷检测图像生成跨域学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。