arXiv:2506.13955stat.MLcs.CR2025-06

用合成异常提升半监督异常检测效果,理论与实证兼备。

Bridging Unsupervised and Semi-Supervised Anomaly Detection: A Theoretically-Grounded and Practical Framework with Synthetic Anomalies

  • 训练时混合真实与合成异常数据,增强模型对异常的识别能力。
  • 在五个基准上均取得显著提升,低密度区域检测性能最优。
  • 适合需要少量标注异常数据的工业场景应用。

异常检测(AD)在网络安全、医疗等领域至关重要。无监督设置下,有效且理论严谨的原则是训练分类器区分正常数据与(合成)异常。本文将该原则扩展至半监督AD,即训练数据包含测试时可能存在的有限标注异常。提出一个兼具理论基础与实证有效的半监督AD框架,训练时融合已知异常与合成异常。为分析半监督AD,首次建立其数学形式化,推广了无监督AD。结果表明,合成异常能(i)改善低密度区域的异常建模,(ii)提供神经网络分类器的最优收敛保证——首个半监督AD的理论结果。在五个不同基准上验证框架有效性,性能持续提升。该优势还扩展至其他基于分类的AD方法,证明合成异常原则具有广泛适用性。

原文摘要 · Abstract (English)

Anomaly detection (AD) is a critical task across domains such as cybersecurity and healthcare. In the unsupervised setting, an effective and theoretically-grounded principle is to train classifiers to distinguish normal data from (synthetic) anomalies. We extend this principle to semi-supervised AD, where training data also include a limited labeled subset of anomalies possibly present in test time. We propose a theoretically-grounded and empirically effective framework for semi-supervised AD that combines known and synthetic anomalies during training. To analyze semi-supervised AD, we introduce the first mathematical formulation of semi-supervised AD, which generalizes unsupervised AD. Here, we show that synthetic anomalies enable (i) better anomaly modeling in low-density regions and (ii) optimal convergence guarantees for neural network classifiers -- the first theoretical result for semi-supervised AD. We empirically validate our framework on five diverse benchmarks, observing consistent performance gains. These improvements also extend beyond our theoretical framework to other classification-based AD methods, validating the generalizability of the synthetic anomaly principle in AD.

异常检测半监督合成数据理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。