首个面向PCB缺陷检测的多模态评测基准,提升工业质检智能化水平
UniPCB: A Unified Vision-Language Benchmark for Open-Ended PCB Quality Inspection
- 构建统一视觉语言数据集,整合多源异构标注数据
- 新模型PCB-GPT在细粒度缺陷定位上性能超竞品一倍以上
- 适合工业质检、多模态大模型研究者参考使用
多模态大语言模型(MLLMs)在通用工业质检中展现潜力,但在印刷电路板(PCB)这类复杂场景下表现不足。PCB检测面临元器件密集、走线结构复杂及微小缺陷模式等挑战,需专业领域知识。然而,缺乏高质量、统一的视觉-语言评测基准,制约了模型评估与进步。为此,我们提出首个开放式的统一视觉-语言基准UniPCB,通过系统化流程从多个来源整合并标准化三类标注场景的数据。同时,我们构建了PCB-GPT,一个基于该流程生成的新指令数据集训练的MLLM,采用模拟人类专家学习过程的渐进式课程。在UniPCB上的评估表明,现有MLLM在特定任务上表现不佳,而PCB-GPT建立新基准,其细粒度缺陷定位性能显著优于最强竞争模型,提升超过一倍。相关指令数据、基准和模型将公开发布,以推动后续研究。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) show promise for general industrial quality inspection, but fall short in complex scenarios, such as Printed Circuit Board (PCB) inspection. PCB inspection poses unique challenges due to densely packed components, complex wiring structures, and subtle defect patterns that require specialized domain expertise. However, a high-quality, unified vision-language benchmark for quantitatively evaluating MLLMs across PCB inspection tasks remains absent, stemming not only from limited data availability but also from fragmented datasets and inconsistent standardization. To fill this gap, we propose UniPCB, the first unified vision-language benchmark for open-ended PCB quality inspection. UniPCB is built via a systematic pipeline that curates and standardizes data from disparate sources across three annotated scenarios. Furthermore, we introduce PCB-GPT, an MLLM trained on a new instruction dataset generated by this pipeline, utilizing a novel progressive curriculum that mimics the learning process of human experts. Evaluations on the UniPCB benchmark show that while existing MLLMs falter on domain-specific tasks, PCB-GPT establishes a new baseline. Notably, it more than doubles the performance on fine-grained defect localization compared to the strongest competitors, with significant advantages in localization and analysis. We will release the instruction data, benchmark, and model to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。