用视觉语言模型实现红外热成像零样本检测缺陷,无需训练数据
Towards Cognitive Defect Analysis in Active Infrared Thermography with Vision-Text Cues
- 基于预训练多模态模型+轻量适配器,实现零样本缺陷理解与定位
- 在25组实际检测序列中,缺陷检测准确率达70%以上,信噪比提升超10dB
- 适合工业无损检测场景,尤其适用于缺乏标注数据的缺陷分析
主动红外热成像(AIRT)正广泛引入人工智能方法,用于高性能碳纤维复合材料(CFRP)的内部缺陷自动化分析。传统AI方法需耗费大量人力物力构建标注数据集进行训练。本文提出一种基于语言引导的新型框架,利用视觉-语言模型(VLMs)实现对CFRP的智能缺陷认知分析。该框架不依赖大规模训练数据,仅通过预训练的多模态编码器结合轻量适配器,即可实现生成式零样本缺陷理解与定位。针对热成像数据与自然图像间的领域差异,提出AIRT-VLM适配器,增强缺陷可见性并对齐热图特征与VLM的语义表示。在包含25组不同冲击能量下缺陷的CFRP检测序列上验证,结果表明,该适配器相比传统降维方法信噪比提升超过10 dB,且实现零样本缺陷检测,交并比达到70%。
原文摘要 · Abstract (English)
Active infrared thermography (AIRT) is currently witnessing a surge of artificial intelligence (AI) methodologies being deployed for automated subsurface defect analysis of high performance carbon fiber-reinforced polymers (CFRP). Deploying AI-based AIRT methodologies for inspecting CFRPs requires the creation of time consuming and expensive datasets of CFRP inspection sequences to train neural networks. To address this challenge, this work introduces a novel language-guided framework for cognitive defect analysis in CFRPs using AIRT and vision-language models (VLMs). Unlike conventional learning-based approaches, the proposed framework does not require developing training datasets for extensive training of defect detectors, instead it relies solely on pretrained multimodal VLM encoders coupled with a lightweight adapter to enable generative zero-shot understanding and localization of subsurface defects. By leveraging pretrained multimodal encoders, the proposed system enables generative zero-shot understanding of thermographic patterns and automatic detection of subsurface defects. Given the domain gap between thermographic data and natural images used to train VLMs, an AIRT-VLM Adapter is proposed to enhance the visibility of defects while aligning the thermographic domain with the learned representations of VLMs. The proposed framework is validated using three representative VLMs; specifically, GroundingDINO, Qwen-VL-Chat, and CogVLM. Validation is performed on 25 CFRP inspection sequences with impacts introduced at different energy levels, reflecting realistic defects encountered in industrial scenarios. Experimental results demonstrate that the AIRT-VLM adapter achieves signal-to-noise ratio (SNR) gains exceeding 10 dB compared with conventional thermographic dimensionality-reduction methods, while enabling zero-shot defect detection with intersection-over-union values reaching 70%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。