构建首个百万级工业缺陷检测多模态基准,支持开闭集场景统一评估。
Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

- 提出新型文本-视觉提示网络,自动生成精准视觉提示。
- 在14类29场景351种缺陷上实现领先性能,推理效率高。
- 适合工业视觉检测、大模型落地研究者参考。
大规模视觉语言模型(LVLM)在自然图像任务中表现优异,但在工业缺陷检测中仍面临两大挑战:一是缺乏覆盖多领域多样缺陷的大规模工业数据集;二是依赖人工输入的提示(点、框、掩码),引入主观噪声且缺乏细粒度文本-视觉交互。为此,我们提出包含超过一百万样本的大型多模态工业开闭集基准MMIOC-1M,涵盖14个超类别、29个工业场景和351个缺陷子类别。据我们所知,这是首个统一支持开集与闭集工业检测的大型基准,可为工业场景下的LVLM提供重要预训练数据。同时,提出改进的文本-视觉提示网络RTVPNet,包含三项创新:(1)专家辅助域投影机制,快速适配通用视觉模型至工业领域;(2)基于能量的稀疏采样策略,无需人工干预自动生成优化视觉提示;(3)双向文本-视觉交互模块,增强跨模态语义对齐与理解。大量实验表明,RTVPNet在MMIOC-1M、LVIS和COCO基准上均达到最优性能,同时保持高效计算。数据集与代码已开源:https://github.com/hellozzk/MMIO。
原文摘要 · Abstract (English)
Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding. To address these challenges, we introduce a Large-Scale Multi-Modal Industrial Open-Closed benchmark (MMIOC-1M) containing over one million samples across $14$ super-categories, $29$ industrial scenes, and $351$ defect subcategories. To our knowledge, MMIOC-1M is the first unified largest benchmark supporting both open-vocabulary and closed-set industrial detection, providing valuable pre-training data for LVLMs in industrial scenarios. Furthermore, we propose a Refined Text-Visual Prompt Network (RTVPNet) that incorporates three key innovations: (1) an expert-assisted domain projection mechanism that enables rapid adaptation of general vision models to industrial domains, (2) an energy-based sparse sampling strategy that automatically generates refined visual prompts without manual intervention, and (3) a bidirectional text-visual interaction module that enhances cross-modal semantic alignment and understanding. Extensive experiments demonstrate that RTVPNet achieves state-of-the-art performance on MMIOC-1M, LVIS, and COCO benchmarks while maintaining computational efficiency. The dataset and code are available at https://github.com/hellozzk/MMIO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。