arXiv:2510.17611cs.CV2025-10被引 5

一个简单框架搞定各种异常检测,性能超越专有模型。

One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection

  • 用五个基础组件构建极简框架,实现统一建模。
  • 多类检测达99.9%和99.3%图像级AUROC,少样本仅需8张正常样本。
  • 支持2D、多视角、多模态等多场景,适合工业、生物等实际应用。

无监督异常检测(UAD)正从专用单类模型转向统一多类模型,但现有方法在多类任务中表现远逊于专用模型。该领域已分化为针对不同场景(多类、3D、少样本等)的专用方法,导致部署困难。本文提出Dinomaly2,首个面向全谱图像无监督异常检测的统一框架,弥合了多类模型的性能差距,并可无缝扩展至多种数据模态与任务设置。基于‘少即是多’理念,我们证明在标准重建框架中,五个简单组件的协同即可实现卓越性能。这种方法论极简性使模型天然适配多样任务而无需修改,表明简约是真正的通用性基础。在12个UAD基准上的实验表明,Dinomaly2在多种模态(2D、多视图、RGB-3D、RGB-IR)、任务设置(单类、多类、推理统一多类、少样本)和应用领域(工业、生物、户外)均表现优越。例如,其多类模型在MVTec-AD和VisA上分别达到99.9%和99.3%的图像级AUROC;在多视图与多模态检测中也达最先进水平,且适应最小。仅用每类8个正常样本,其性能即超越以往全样本模型,在MVTec-AD和VisA上分别达到98.7%和97.4%的图像级AUROC。极简设计、计算可扩展性与普适性使Dinomaly2成为真实世界异常检测的统一解决方案。

原文摘要 · Abstract (English)

Unsupervised anomaly detection (UAD) has evolved from building specialized single-class models to unified multi-class models, yet existing multi-class models significantly underperform the most advanced one-for-one counterparts. Moreover, the field has fragmented into specialized methods tailored to specific scenarios (multi-class, 3D, few-shot, etc.), creating deployment barriers and highlighting the need for a unified solution. In this paper, we present Dinomaly2, the first unified framework for full-spectrum image UAD, which bridges the performance gap in multi-class models while seamlessly extending across diverse data modalities and task settings. Guided by the "less is more" philosophy, we demonstrate that the orchestration of five simple element achieves superior performance in a standard reconstruction-based framework. This methodological minimalism enables natural extension across diverse tasks without modification, establishing that simplicity is the foundation of true universality. Extensive experiments on 12 UAD benchmarks demonstrate Dinomaly2's full-spectrum superiority across multiple modalities (2D, multi-view, RGB-3D, RGB-IR), task settings (single-class, multi-class, inference-unified multi-class, few-shot) and application domains (industrial, biological, outdoor). For example, our multi-class model achieves unprecedented 99.9% and 99.3% image-level (I-) AUROC on MVTec-AD and VisA respectively. For multi-view and multi-modal inspection, Dinomaly2 demonstrates state-of-the-art performance with minimum adaptations. Moreover, using only 8 normal examples per class, our method surpasses previous full-shot models, achieving 98.7% and 97.4% I-AUROC on MVTec-AD and VisA. The combination of minimalistic design, computational scalability, and universal applicability positions Dinomaly2 as a unified solution for the full spectrum of real-world anomaly detection applications.

异常检测统一框架少样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。