arXiv:2510.06505cs.LGcs.AI2025-10中稿 · TMLR被引 3

用无标签数据中的梯度中位数,自动发现异常样本并提升检测效果。

Medix: Out-of-Distribution Detection from Unlabeled Wild Data via Robust Gradient Statistics

  • 基于梯度中位数统计,从无标签野数据中识别潜在异常点。
  • 在开放世界设置下,各项指标均优于现有方法,误差率更低。
  • 适合需要高鲁棒性的实际部署场景,尤其适用于无标注数据环境。

分布外(OOD)检测对于保障机器学习系统在真实应用中的鲁棒性至关重要。近期方法尝试利用无标签数据,展现出增强OOD检测能力的潜力。然而,由于无标签野数据中混合了分布内(InD)和分布外样本,有效利用这些数据仍具挑战性。缺乏明确的分布外样本使训练最优的OOD分类器变得复杂。本文提出Medix,一种新框架,通过中位数为基础的鲁棒梯度统计,从无标签数据中识别潜在异常点。使用中位数是因为其对噪声和异常值具有强鲁棒性,能稳定估计中心趋势。利用这些识别出的异常点与有标签的InD数据,训练出更鲁棒的OOD分类器。理论分析表明,Medix具有较低的误差界。实验结果进一步验证了该方法的有效性,在开放世界设置下全面超越现有方法。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection plays a crucial role in ensuring the robustness of machine learning systems deployed in real-world applications. Recent approaches have explored the use of unlabeled data, showing potential for enhancing OOD detection capabilities. However, effectively utilizing unlabeled in-the-wild data remains challenging due to the mixed nature of both in-distribution (InD) and OOD samples. The lack of a distinct set of OOD samples complicates the task of training an optimal OOD classifier. In this work, we introduce Medix, a novel framework designed to identify potential outliers from unlabeled data using the median-based robust gradient statistics. We use the median because it provides a stable estimate of the central tendency, as an OOD detection mechanism, due to its robustness against noise and outliers. Using these identified outliers, along with labeled InD data, we train a robust OOD classifier. From a theoretical perspective, we derive error bounds that demonstrate Medix achieves a low error rate. Empirical results further substantiate our claims, as Medix outperforms existing methods across the board in open-world settings.

OOD检测无监督学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。