arXiv:2501.08821cs.LG2025-01被引 1

揭示了分布外检测在何种条件下可学习,突破传统理论悲观结论。

A Closer Look at the Learnability of Out-of-Distribution (OOD) Detection

  • 区分均匀与非均匀可学习性,重构理论框架
  • 证明多个此前被认为不可学习的情况实际可解
  • 提供可落地的算法与样本复杂度分析,适合研究者参考

机器学习模型在实际部署时经常遇到分布外(OOD)数据,通常依赖OOD检测来识别此类样本。尽管实践中表现良好,但现有理论对OOD检测的分析普遍悲观。本文从PAC学习理论出发,区分均匀可学习性与非均匀可学习性,刻画了在何种条件下OOD检测具有可学习性。研究表明,在若干情形下,非均匀可学习性将一系列负面结果转为正面。对于所有可学习的情形,本文均提供了具体的学习算法及样本复杂度分析。

原文摘要 · Abstract (English)

Machine learning algorithms often encounter different or "out-of-distribution" (OOD) data at deployment time, and OOD detection is frequently employed to detect these examples. While it works reasonably well in practice, existing theoretical results on OOD detection are highly pessimistic. In this work, we take a closer look at this problem, and make a distinction between uniform and non-uniform learnability, following PAC learning theory. We characterize under what conditions OOD detection is uniformly and non-uniformly learnable, and we show that in several cases, non-uniform learnability turns a number of negative results into positive. In all cases where OOD detection is learnable, we provide concrete learning algorithms and a sample-complexity analysis.

OOD检测理论分析学习能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。