通过提升辅助异常数据多样性,显著增强模型对未知异常的检测能力。
Out-Of-Distribution Detection with Diversification (Provably)
- 用混合法增强训练时辅助异常数据的多样性
- 在多个主流和挑战性基准上表现更优
- 提供理论保证,适合部署可靠性要求高的场景
分布外(OOD)检测对保障机器学习模型可靠部署至关重要。近期方法依赖易获取的辅助异常数据(如网络或其它数据集数据)进行训练,但我们实验发现这些方法在面对未知异常数据时仍难以泛化,原因是所收集的辅助异常数据多样性有限。从泛化视角深入分析后,我们证明:更丰富的辅助异常数据对提升检测能力至关重要。然而,实际中获取足够多样化的异常数据成本高昂。为此,我们提出一种简单且实用的方法——多样性诱导混合法(diverseMix),以高效方式增强训练时辅助异常数据的多样性,并提供理论保证。大量实验表明,diverseMix在常用及最新大型基准上均取得优异性能,进一步验证了辅助异常数据多样性的关键作用。
原文摘要 · Abstract (English)
Out-of-distribution (OOD) detection is crucial for ensuring reliable deployment of machine learning models. Recent advancements focus on utilizing easily accessible auxiliary outliers (e.g., data from the web or other datasets) in training. However, we experimentally reveal that these methods still struggle to generalize their detection capabilities to unknown OOD data, due to the limited diversity of the auxiliary outliers collected. Therefore, we thoroughly examine this problem from the generalization perspective and demonstrate that a more diverse set of auxiliary outliers is essential for enhancing the detection capabilities. However, in practice, it is difficult and costly to collect sufficiently diverse auxiliary outlier data. Therefore, we propose a simple yet practical approach with a theoretical guarantee, termed Diversity-induced Mixup for OOD detection (diverseMix), which enhances the diversity of auxiliary outlier set for training in an efficient way. Extensive experiments show that diverseMix achieves superior performance on commonly used and recent challenging large-scale benchmarks, which further confirm the importance of the diversity of auxiliary outliers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。