用少样本数据提升管道缺陷分割精度,兼顾性能与效率。
AI-Based Culvert-Sewer Inspection
- 结合数据增强与标签注入,提升小样本分割效果
- 新模型FORTRESS参数量少、计算成本低,表现优于现有方法
- 采用少样本学习框架,适合标注数据稀缺的工程场景
涵洞与污水管是排水系统的关键组成部分,其失效可能对公共安全和环境造成严重威胁。本文研究如何在标注数据有限的情况下提升涵洞与污水管缺陷的自动化分割能力。由于该领域数据收集与标注繁琐且需专业知识,大规模结构缺陷数据集难以获取。为此,本文提出三种方法以显著提升缺陷分割性能并应对数据稀缺问题,可通过增强训练数据或调整模型架构实现。首先,评估了包括传统数据增强和动态标签注入在内的预处理策略,显著提升了分割性能,交并比(IoU)与F1分数均得到提高。其次,提出FORTRESS新架构,融合深度可分离卷积、自适应柯尔莫戈罗夫-阿诺德网络(KAN)及多尺度注意力机制,在涵洞污水管缺陷数据集上达到当前最优表现,同时大幅降低可训练参数数量与计算开销。最后,探索少样本语义分割在缺陷检测中的应用,通过双向原型网络结合注意力机制,获得更丰富的特征表示,在各项评估指标上取得满意结果。
原文摘要 · Abstract (English)
Culverts and sewer pipes are critical components of drainage systems, and their failure can lead to serious risks to public safety and the environment. In this thesis, we explore methods to improve automated defect segmentation in culverts and sewer pipes. Collecting and annotating data in this field is cumbersome and requires domain knowledge. Having a large dataset for structural defect detection is therefore not feasible. Our proposed methods are tested under conditions with limited annotated data to demonstrate applicability to real-world scenarios. Overall, this thesis proposes three methods to significantly enhance defect segmentation and handle data scarcity. This can be addressed either by enhancing the training data or by adjusting a models architecture. First, we evaluate preprocessing strategies, including traditional data augmentation and dynamic label injection. These techniques significantly improve segmentation performance, increasing both Intersection over Union (IoU) and F1 score. Second, we introduce FORTRESS, a novel architecture that combines depthwise separable convolutions, adaptive Kolmogorov-Arnold Networks (KAN), and multi-scale attention mechanisms. FORTRESS achieves state-of-the-art performance on the culvert sewer pipe defect dataset, while significantly reducing the number of trainable parameters, as well as its computational cost. Finally, we investigate few-shot semantic segmentation and its applicability to defect detection. Few-shot learning aims to train models with only limited data available. By employing a bidirectional prototypical network with attention mechanisms, the model achieves richer feature representations and achieves satisfactory results across evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。