一个模型搞定工业缺陷检测、解释和编辑,三者统一评估。
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
- 用区域特征注入让视觉语言模型精准定位缺陷
- 在59,916张图上实现缺陷分割、解释与生成统一
- 适合需要缺陷理解与可控编辑的工业质检场景
真实工业检测不仅需定位缺陷,还需用自然语言解释并生成可控的缺陷修改。现有方法无法在统一框架与评估协议下同时支持这三项能力。我们提出IAD-Unify,一种双编码器统一框架:冻结的DINOv2区域专家通过轻量级标记注入,将精确异常证据提供给共享的Qwen3.5-4B视觉语言主干,联合实现异常分割、基于区域的理解与掩码引导生成。为实现统一评估,我们进一步构建Anomaly-56K,一个涵盖59,916张图像、24类、104种缺陷变体的综合性多任务工业异常检测评估平台。受控消融实验得出四点发现:(i) 区域定位是理解的关键机制,移除后定位准确率下降超76个百分点;(ii) 预测区域性能接近理想情况,证实部署可行性;(iii) 基于区域的生成在全图保真度与掩码区域感知质量上表现最佳;(iv) 预初始化联合训练在生成代价几乎不变(-0.16 dB)的前提下提升理解能力。IAD-Unify在MMAD基准上也表现出色,包括训练中未见类别,证明其强大的跨类别泛化能力。
原文摘要 · Abstract (English)
Real-world industrial inspection requires not only localizing defects, but also explaining them in natural language and generating controlled defect edits. However, existing approaches fail to jointly support all three capabilities within a unified framework and evaluation protocol. We propose IAD-Unify, a dual-encoder unified framework in which a frozen DINOv2-based region expert supplies precise anomaly evidence to a shared Qwen3.5-4B vision-language backbone via lightweight token injection, jointly enabling anomaly segmentation, region-grounded understanding, and mask-guided generation. To enable unified evaluation, we further construct Anomaly-56K, a comprehensive unified multi-task IAD evaluation platform, spanning 59,916 images across 24 categories and 104 defect variants. Controlled ablations yield four findings: (i) region grounding is the decisive mechanism for understanding, removing it degrades location accuracy by >76 pp; (ii) predicted-region performance closely matches oracle, confirming deployment viability; (iii) region-grounded generation achieves the best full-image fidelity and masked-region perceptual quality; and (iv) pre-initialized joint training improves understanding at negligible generation cost (-0.16 dB). IAD-Unify further achieves strong performance on the MMAD benchmark, including categories unseen during training, demonstrating robust cross-category generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。