arXiv:2506.14194cs.LG2025-06被引 4

用信息论设计可解释的异常检测特征,提升模型泛化能力。

An Information-Theoretic Framework for Feature Construction in Out-of-Distribution Detection

  • 基于信息瓶颈与KL散度构建新损失函数,分离分布特征。
  • 提出的新特征在多种分布外场景下性能超越现有方法。
  • 框架可生成多种可解释特征,适合安全敏感场景使用。

我们提出一种用于神经网络分布外(OOD)检测特征构造的信息论框架。通过引入一种新颖的信息论损失函数,该函数包含两项:第一项基于KL散度以分离分布内(ID)与分布外(OOD)特征分布;第二项为信息瓶颈,促使压缩特征保留关键的OOD信息。我们制定了变分优化过程以求解该损失,获得有效的OOD特征。在对OOD分布的合理假设下,可推导出已有特征的形状函数性质。进一步地,我们的理论预测了一种新型形状函数,在包括语义差异和广泛协变量偏移在内的多个基准测试中表现优于现有方法。该理论提供了一个通用框架,可构建具有明确可解释性的多样化新特征。

原文摘要 · Abstract (English)

We present a theory for the construction of out-of-distribution (OOD) detection features for neural networks. We introduce random features for OOD through a novel information-theoretic loss functional consisting of two terms, the first based on the KL divergence separates resulting in-distribution (ID) and OOD feature distributions and the second term is the Information Bottleneck, which favors compressed features that retain the OOD information. We formulate a variational procedure to optimize the loss and obtain OOD features. Based on assumptions on OOD distributions, one can recover properties of existing OOD features, i.e., shaping functions. Furthermore, we show that our theory can predict a new shaping function that out-performs existing ones on OOD benchmarks, including both semantic and a wide range of covariate shifts. Our theory provides a general framework for constructing a variety of new features with clear explainability.

OOD检测信息论可解释性特征构造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。