让视觉语言模型读懂红外工业图像,实现无标签零件检测。
Vision-Language Models for Infrared Industrial Sensing in Additive Manufacturing Scene Description
- 将红外图像转为类熔岩色彩表示,适配现有视觉语言模型。
- 在3D打印场景中零样本检测工件存在,准确率高无需重训练。
- 适合工业质检、无人监控等需热成像的场景快速部署。
许多制造环境光照不足或处于封闭设备内,传统视觉系统难奏效。红外相机在此类场景具优势。同时,监督学习依赖大量标注数据,零样本学习更实用。近年视觉-语言基础模型(VLMs)通过图像-文本对实现零样本预测,但现有模型仅训练于可见光数据,无法理解红外图像。本文提出VLM-IRIS框架,通过将FLIR Boson传感器采集的红外图像预处理为适配CLIP编码器的RGB格式,使模型可处理热成像数据。在3D打印机工作台上,利用基板与工件间的温差进行零样本工件存在性检测。VLM-IRIS采用熔岩色表示转换与质心提示集成策略,结合CLIP ViT-B/32编码器,在不重新训练模型的前提下实现高精度识别。结果表明,该方法可有效拓展视觉语言模型至热成像应用,实现免标签工业监测。
原文摘要 · Abstract (English)
Many manufacturing environments operate in low-light conditions or within enclosed machines where conventional vision systems struggle. Infrared cameras provide complementary advantages in such environments. Simultaneously, supervised AI systems require large labeled datasets, which makes zero-shot learning frameworks more practical for applications including infrared cameras. Recent advances in vision-language foundation models (VLMs) offer a new path in zero-shot predictions from paired image-text representations. However, current VLMs cannot understand infrared camera data since they are trained on RGB data. This work introduces VLM-IRIS (Vision-Language Models for InfraRed Industrial Sensing), a zero-shot framework that adapts VLMs to infrared data by preprocessing infrared images captured by a FLIR Boson sensor into RGB-compatible inputs suitable for CLIP-based encoders. We demonstrate zero-shot workpiece presence detection on a 3D printer bed where temperature differences between the build plate and workpieces make the task well-suited for thermal imaging. VLM-IRIS converts the infrared images to magma representation and applies centroid prompt ensembling with a CLIP ViT-B/32 encoder to achieve high accuracy on infrared images without any model retraining. These findings demonstrate that the proposed improvements to VLMs can be effectively extended to thermal applications for label-free monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。