arXiv:2606.08033cs.CVcs.LG2026-06

用合成数据辅助真实数据,提升砖石结构裂缝检测精度。

Balancing Real and Synthetic Data for CNN-based Masonry Crack Detection

论文配图:Balancing Real and Synthetic Data for CNN-based Masonry Crack Detection
图 1 · 摘自论文原文
  • 用可控方式生成裂缝图像,与真实数据混合训练
  • 20%真实+80%合成数据达76%准确率,超纯真实数据
  • 适合缺乏真实数据的建筑健康监测场景

裂缝是建筑健康的关键指标,早期识别对防止损害至关重要。深度学习(特别是卷积神经网络,CNN)为自动化裂缝检测提供了可扩展解决方案,但其性能依赖于大规模且多样化的数据集,这对复杂表面如砖石结构尤为困难。真实数据收集耗时,公开数据集又常不足。为此,本文探索生成合成裂缝数据以补充真实数据,提升训练效果。真实数据来自意大利博洛尼亚及其周边建筑的砖石裂缝图像,合成数据通过裂缝叠加工具在背景图上可控地添加裂缝。基于真实数据训练多个深度学习架构,筛选出表现最佳的InceptionV4模型用于后续实验。在InceptionV4上测试了六种真实与合成数据比例组合,评估使用真实测试集的F1-score和平均交并比(mIoU)。结果表明,仅需20%真实数据加80%合成数据,即可达到与纯真实数据训练相当的效果;其中20/80(合成/真实)配置取得76% F1-score和80% mIoU,优于纯真实数据训练。该方法证明了合成数据能显著减少采集成本,同时提升检测精度。

原文摘要 · Abstract (English)

Cracks are a critical indicator of building health, and early stage identification is fundamental to prevent harmful damages. Advances in deep learning (DL), particularly convolutional neural networks (CNNs), have enabled scalable solutions for automated crack detection. However, CNN performance strongly depends on the availability of large and diverse datasets, which is particularly challenging for complex surfaces such as masonry. Collecting sufficient real data is time-consuming, while publicly available datasets may not be adequate. To address this limitation, we explored generating synthetic crack data, which complements real data and improves training effectiveness. The real dataset consists of masonry crack images collected from buildings in Bologna and surrounding areas. In contrast, the synthetic dataset was generated using a crack overlay tool that adds cracks to background images in a controlled orientation and placement. The real dataset was used to train several DL architectures, to identify the best-performing model (InceptionV4) employed for experiments with generated data. Six training scenarios were tested in InceptionV4 by varying the ratio of real and synthetic data, with evaluation performed on a test set composed of real images using the F1-score and mean Intersection over Union (mIoU) metrics. Results show that training on synthetic data plus a modest addition of 20% real data achieves results comparable to training on real data only. Moreover, the 20/80 scenario (synthetic/real) achieved an 76% F1-score and 80% mean IoU, outperforming the real-only case. As can be seen, the method demonstrates the potential of synthetic data to reduce collection efforts while enhancing crack detection accuracy.

裂缝检测合成数据CNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。