用视觉语言模型细调,精准识别钢铁废料中的异物
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
- 用多尺度图像编码器和文本提示对齐正负样本
- 通过多分类监督微调实现细粒度异常检测
- 适合钢铁回收场景的自动化质检系统
回收钢铁废料可降低钢铁行业的二氧化碳(CO2)排放。然而,钢铁回收中的一大挑战是混入非钢杂质。为解决此问题,我们提出基于视觉语言模型的异常检测方法,通过有监督微调使模型能有效处理特定类别的异常对象。该模型可在钢铁废料中实现细粒度异常检测。具体而言,我们对配备多尺度机制的图像编码器进行微调,并使用与正常及异常图像对齐的文本提示。微调过程采用多分类作为监督信号训练这些模块。
原文摘要 · Abstract (English)
Recycling steel scrap can reduce carbon dioxide (CO2) emissions from the steel industry. However, a significant challenge in steel scrap recycling is the inclusion of impurities other than steel. To address this issue, we propose vision-language-model-based anomaly detection where a model is finetuned in a supervised manner, enabling it to handle niche objects effectively. This model enables automated detection of anomalies at a fine-grained level within steel scrap. Specifically, we finetune the image encoder, equipped with multi-scale mechanism and text prompts aligned with both normal and anomaly images. The finetuning process trains these modules using a multiclass classification as the supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。