用重叠指数构建轻量级可信度分数,高效识别分布外样本
Out-of-Distribution Detection with Overlap Index
- 基于重叠指数设计非参数化可信度函数,无需调参且易解释
- 在多数据集上达到顶尖检测精度,计算与内存开销更低
- 对分布微小变化不敏感,适合高维数据和模型可靠性评估
分布外(OOD)检测对机器学习模型在开放世界中的部署至关重要。现有方法虽能有效识别明显偏离正常分布的样本,但常面临权衡:深度方法计算成本高、需调参且可解释性差;传统方法在大规模高维数据上准确率较低。为此,我们提出一种基于重叠指数(OI)的新型高效OOD检测方法,通过构造非参数化、轻量且易解释的可信度分数来评估输入与已知内分布样本的一致性。大量实验证明,该方法在多种数据集上检测精度媲美前沿模型,同时显著降低计算与内存开销。此外,所提可信度函数继承了重叠指数的优良性质,如对小分布变化不敏感、对Huber ε-污染鲁棒,是估计重叠指数及模型准确性的通用工具。
原文摘要 · Abstract (English)
Out-of-distribution (OOD) detection is crucial for the deployment of machine learning models in the open world. While existing OOD detectors are effective in identifying OOD samples that deviate significantly from in-distribution (ID) data, they often come with trade-offs. For instance, deep OOD detectors usually suffer from high computational costs, require tuning hyperparameters, and have limited interpretability, whereas traditional OOD detectors may have a low accuracy on large high-dimensional datasets. To address these limitations, we propose a novel effective OOD detection approach that employs an overlap index (OI)-based confidence score function to evaluate the likelihood of a given input belonging to the same distribution as the available ID samples. The proposed OI-based confidence score function is non-parametric, lightweight, and easy to interpret, hence providing strong flexibility and generality. Extensive empirical evaluations indicate that our OI-based OOD detector is competitive with state-of-the-art OOD detectors in terms of detection accuracy on a wide range of datasets while requiring less computation and memory costs. Lastly, we show that the proposed OI-based confidence score function inherits nice properties from OI (e.g., insensitivity to small distributional variations and robustness against Huber $ε$-contamination) and is a versatile tool for estimating OI and model accuracy in specific contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。