通过可解模型揭示异常检测中的类别不平衡机制
Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model
- 基于教师-学生感知机的精确解,建立理论分析框架
- 发现最优训练不平衡度非50%,受数据量和噪声影响
- 揭示小噪声与大噪声下的性能拐点,指导实际调参
类别不平衡(CI)是机器学习中的长期难题,导致训练变慢、性能下降。尽管存在一些经验性解决方法,但由于缺乏统一理论,难以判断何时何法最有效。本文聚焦异常检测这一常见场景,基于对教师-学生感知机模型的精确解(通过复制理论),建立了分析、解释和应对类别不平衡的理论框架。该框架可区分内在不平衡、训练集不平衡与测试集不平衡。分析表明,最优训练不平衡度通常不等于50%,且其值依赖于内在不平衡程度、数据丰富度以及学习过程中的噪声水平。此外,存在一个从低噪声训练(性能与噪声无关)到高噪声训练(性能随噪声迅速下降)的相变点。研究结果挑战了部分关于类别不平衡的传统认知,并提供了可操作的实践指导。
原文摘要 · Abstract (English)
Class imbalance (CI) is a longstanding problem in machine learning, slowing down training and reducing performances. Although empirical remedies exist, it is often unclear which ones work best and when, due to the lack of an overarching theory. We address a common case of imbalance, that of anomaly (or outlier) detection. We provide a theoretical framework to analyze, interpret and address CI. It is based on an exact solution of the teacher-student perceptron model, through replica theory. Within this framework, one can distinguish several sources of CI: either intrinsic, train or test imbalance. Our analysis reveals that the optimal train imbalance is generally different from 50%, with a non trivial dependence on the intrinsic imbalance, the abundance of data and on the noise in the learning. Moreover, there is a crossover between a small noise training regime where results are independent of the noise level to a high noise regime where performances quickly degrade with noise. Our results challenge some of the conventional wisdom on CI and offer practical guidelines to address it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。