arXiv:2505.05273stat.MLcs.IT2025-05被引 1

用巴塔查里亚散度设计更温和的拒绝策略,提升模型可靠性。

A Connection Between Learning to Reject and Bhattacharyya Divergences

  • 基于输入与标签的联合理想分布构建拒绝机制
  • 发现联合分布下的拒绝等价于阈值化偏斜巴塔查里亚散度
  • 相比传统方法更少拒绝,适合对误拒敏感的应用

学习拒绝是一种允许模型在不确定时放弃预测的学习范式。以往方法通过学习输入域的理想边缘分布,并与真实分布比较密度比来实现。本文提出考虑输入与标签的联合理想分布,建立拒绝与统计散度阈值之间的联系。研究发现,当使用一种变体对数损失时,基于联合分布的拒绝器等价于对类别概率间偏斜巴塔查里亚散度进行阈值化。这与仅考虑边缘分布的典型最优拒绝准则——乔规则(Chow's Rule)——对应于KL散度阈值化形成对比。总体而言,基于巴塔查里亚散度的拒绝策略比乔规则更为温和。

原文摘要 · Abstract (English)

Learning to reject provide a learning paradigm which allows for our models to abstain from making predictions. One way to learn the rejector is to learn an ideal marginal distribution (w.r.t. the input domain) - which characterizes a hypothetical best marginal distribution - and compares it to the true marginal distribution via a density ratio. In this paper, we consider learning a joint ideal distribution over both inputs and labels; and develop a link between rejection and thresholding different statistical divergences. We further find that when one considers a variant of the log-loss, the rejector obtained by considering the joint ideal distribution corresponds to the thresholding of the skewed Bhattacharyya divergence between class-probabilities. This is in contrast to the marginal case - that is equivalent to a typical characterization of optimal rejection, Chow's Rule - which corresponds to a thresholding of the Kullback-Leibler divergence. In general, we find that rejecting via a Bhattacharyya divergence is less aggressive than Chow's Rule.

拒绝学习散度度量概率校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。