arXiv:2606.07965cs.AI2026-06AAAI被引 9

为工业缺陷检测构建新数据集并提出更精准的零样本提示方法

Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline

论文配图:Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline
图 1 · 摘自论文原文
  • 设计基于Mobile-SAM的专家引导域适应机制,提升模型在工业场景泛化能力
  • 自动从图像生成视觉提示,交互式理解图文内容,零样本与闭样本准确率分别达42.2%和24.7%
  • 发布首个大规模工业多场景预训练数据集MMIO,含8万+样本,覆盖18个子类别

大型视觉语言模型(LVLM)在自然场景中表现优异,但在工业场景应用面临挑战。现有方法依赖用户输入提示进行物体分割,常因包含无关像素导致性能下降;且工业数据稀缺限制了其应用。本文提出一个开放的工业数据集和一种改进的文本-视觉提示方法(RTVP)。首先构建了多模态工业开放数据集(MMIO),包含超过8万张样本,涵盖6大类、18个子类,是首个面向工业零样本学习的大规模多场景预训练数据集。基于此,提出专门用于工业零样本任务的RTVP:一是设计专家引导的大模型领域自适应机制,结合Mobile-SAM实现工业场景下的强泛化能力;二是自动从图像生成视觉提示,并建模文本-视觉提示交互,增强图文理解。RTVP在MMIO的零样本与闭样本测试中分别达到42.2%和24.7%的AP,达到当前最优水平。

原文摘要 · Abstract (English)

Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks. However, the significant differences between industrial and natural scenes make applying LVLMs challenging. Existing LVLMs rely on user-provided prompts to segment objects. This often leads to suboptimal performance due to the inclusion of irrelevant pixels. In addition, the scarcity of data also makes the application of LVLMs in industrial scenarios remain unexplored. To fill this gap, this paper proposes an open industrial dataset and a Refined Text-Visual Prompt (RTVP) for zero-shot industrial defect detection. First, this paper constructs the Multi-Modal Industrial Open Dataset (MMIO) containing 80K+ samples. MMIO contains diverse industrial categories, including 6 super categories and 18 subcategories. MMIO is the first large-scale multi-scenes pre-training dataset for industrial zero-shot learning, and provides valuable training data for open models in future industrial scenarios. Based on MMIO, this paper provides a RTVP specifically for industrial zero-shot tasks. RTVP has two significant advantages: First, this paper designs an expert-guided large model domain adaptation mechanism and designs an industrial zero-shot method based on Mobile-SAM, which enhances the generalization ability of large models in industrial scenarios. Second, RTVP automatically generates visual prompts directly from images and considers text-visual prompt interactions ignored by previous LVLM, improving visual and textual content understanding. RTVP achieves SOTA with 42.2% and 24.7% AP in zero-shot and closed scenes of MMIO.

工业缺陷检测零样本学习视觉语言模型多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。