arXiv:2512.24160cs.CV2025-12被引 3

首个百万级工业缺陷多模态数据集,助力智能质检模型高效泛化。

Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset

  • 构建百万级图文对数据集,覆盖60类材料400+缺陷类型
  • 基于该数据集训练的扩散模型仅用5%数据即达专家模型性能
  • 适合工业质检、缺陷生成与跨场景迁移研究者使用

我们提出IMDD-1M,首个大规模工业多模态缺陷数据集,包含100万组对齐的图像-文本配对,涵盖超过60种材料类别和400多种缺陷类型。每条数据均配有专家验证的标注及细粒度文本描述,涵盖缺陷位置、严重程度与上下文属性。该数据集支持分类、分割、检索、描述生成与生成建模等多种应用。基于IMDD-1M,我们从零训练了一个专为工业场景设计的扩散型视觉语言基础模型。该模型可通过轻量微调高效适配特定领域,仅需任务数据的5%即可达到专用专家模型性能,凸显数据高效基础模型适应在工业检测与生成中的潜力,推动可扩展、领域自适应、知识驱动的制造智能化发展。更多信息请访问:https://ninaneon.github.io/projectpage/

原文摘要 · Abstract (English)

We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution real-world defects spanning over 60 material categories and more than 400 defect types, each accompanied by expert-verified annotations and fine-grained textual descriptions detailing defect location, severity, and contextual attributes. This dataset enables a wide spectrum of applications, including classification, segmentation, retrieval, captioning, and generative modeling. Building upon IMDD-1M, we train a diffusion-based vision-language foundation model from scratch, specifically tailored for industrial scenarios. The model serves as a generalizable foundation that can be efficiently adapted to specialized domains through lightweight fine-tuning. With less than 5% of the task-specific data required by dedicated expert models, it achieves comparable performance, highlighting the potential of data-efficient foundation model adaptation for industrial inspection and generation, paving the way for scalable, domain-adaptive, and knowledge-grounded manufacturing intelligence. Additional details and resources can be found in this URL: https://ninaneon.github.io/projectpage/

工业质检多模态扩散模型数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。