用生成模型和视觉变压器提升道路病害自动识别精度
Automated Road Distress Detection Using Vision Transformersand Generative Adversarial Networks
- 用GAN生成合成数据增强训练,提升模型泛化能力
- 基于Transformer的MaskFormer在两个指标上优于传统CNN
- 适合智能交通与基础设施监测方向的研究者参考
美国土木工程师协会将美国基础设施状况评为C级,道路系统更仅为D级。道路对区域经济至关重要,但其管理维护仍依赖过时的手动或激光检测方法,成本高且耗时。随着自动驾驶车辆实时视觉数据的增多,利用计算机视觉技术实现道路状态智能监控成为可能。本研究探索先进计算机视觉方法在道路病害分割中的应用。首先评估生成对抗网络(GAN)生成的合成数据对模型训练的有效性;接着使用卷积神经网络(CNN)进行病害分割,随后考察基于变换器的模型MaskFormer。结果表明,使用GAN生成的数据可提升模型性能,且MaskFormer在mAP50和IoU两项指标上均优于CNN模型。
原文摘要 · Abstract (English)
The American Society of Civil Engineers has graded Americas infrastructure condition as a C, with the road system receiving a dismal D. Roads are vital to regional economic viability, yet their management, maintenance, and repair processes remain inefficient, relying on outdated manual or laser-based inspection methods that are both costly and time-consuming. With the increasing availability of real-time visual data from autonomous vehicles, there is an opportunity to apply computer vision (CV) methods for advanced road monitoring, providing insights to guide infrastructure rehabilitation efforts. This project explores the use of state-of-the-art CV techniques for road distress segmentation. It begins by evaluating synthetic data generated with Generative Adversarial Networks (GANs) to assess its usefulness for model training. The study then applies Convolutional Neural Networks (CNNs) for road distress segmentation and subsequently examines the transformer-based model MaskFormer. Results show that GAN-generated data improves model performance and that MaskFormer outperforms the CNN model in two metrics: mAP50 and IoU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。