用视觉语言模型提升无人机巡检缺陷检测准确率
Transmission Line Defect Detection Based on UAV Patrol Images and Vision-language Pretraining
- 结合视觉与语言预训练,增强图像编码器表征能力
- 在真实无人机图像上实现缺陷检测准确率显著提升
- 适合电力巡检、低资源缺陷识别场景使用
无人机巡检因成本低已成为输电线路监测的主要方式。但受拍摄距离和角度限制,无人机图像常缺乏足够的缺陷视觉信息,影响检测精度。本文提出基于视觉语言预训练的输电线路检测方法(VLP-TL)及渐进式迁移策略(PTS)。VLP-TL设计了两个面向输电线路场景的新型预训练任务,利用视觉与语言信息联合训练图像编码器,使其获得更丰富的知识表征。将预训练编码器作为缺陷检测器主干网络,有效缓解图像信息不足问题。此外,PTS通过渐进式桥接预训练与下游检测任务间的差距,进一步提升迁移性能。实验表明,该方法通过融合多模态信息,显著提升了缺陷检测准确率,克服了无人机图像中缺陷相关视觉信息不足的局限。
原文摘要 · Abstract (English)
Unmanned aerial vehicle (UAV) patrol inspection has emerged as a predominant approach in transmission line monitoring owing to its cost-effectiveness. Detecting defects in transmission lines is a critical task during UAV patrol inspection. However, due to imaging distance and shooting angles, UAV patrol images often suffer from insufficient defect-related visual information, which has an adverse effect on detection accuracy. In this article, we propose a novel method for detecting defects in UAV patrol images, which is based on vision-language pretraining for transmission line (VLP-TL) and a progressive transfer strategy (PTS). Specifically, VLP-TL contains two novel pretraining tasks tailored for the transmission line scenario, aimimg at pretraining an image encoder with abundant knowledge acquired from both visual and linguistic information. Transferring the pretrained image encoder to the defect detector as its backbone can effectively alleviate the insufficient visual information problem. In addition, the PTS further improves transfer performance by progressively bridging the gap between pretraining and downstream defection detection. Experimental results demonstrate that the proposed method significantly improves defect detection accuracy by jointly utilizing multimodal information, overcoming the limitations of insufficient defect-related visual information provided by UAV patrol images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。