arXiv:2507.23136cs.LG2025-07被引 2

研究分类模型预测的不确定性,发现某些群体预测更易变动,影响安全决策。

Observational Multiplicity

  • 用‘后悔值’衡量模型预测随标签变化的波动性。
  • 发现部分数据群体的预测后悔值更高,易产生不一致结果。
  • 可指导模型拒绝预测或收集新数据,提升实际应用安全性。

许多预测任务可存在多个表现相近的模型,这种现象在概率分类中可能引发任意性,导致不同模型对同一个体给出冲突预测,影响可解释性和安全性。本文研究了由‘观测多重性’引发的任意性,即当学习分类器预测概率 $p_i \in [0,1]$ 但仅获得二值观测 $y_i \in \{0,1\}$ 时的情况。我们提出通过‘后悔值’评估个体预测的任意程度,定义了一种衡量概率分类中预测随训练标签变化敏感性的指标。提出一种通用方法估计该后悔值,并实证显示某些群体的后悔值更高。进一步展示通过估计后悔值可实现模型拒判与数据收集,从而增强现实应用的安全性。

原文摘要 · Abstract (English)

Many prediction tasks can admit multiple models that can perform almost equally well. This phenomenon can can undermine interpretability and safety when competing models assign conflicting predictions to individuals. In this work, we study how arbitrariness can arise in probabilistic classification tasks as a result of an effect that we call \emph{observational multiplicity}. We discuss how this effect arises in a broad class of practical applications where we learn a classifier to predict probabilities $p_i \in [0,1]$ but are given a dataset of observations $y_i \in \{0,1\}$. We propose to evaluate the arbitrariness of individual probability predictions through the lens of \emph{regret}. We introduce a measure of regret for probabilistic classification tasks, which measures how the predictions of a model could change as a result of different training labels change. We present a general-purpose method to estimate the regret in a probabilistic classification task. We use our measure to show that regret is higher for certain groups in the dataset and discuss potential applications of regret. We demonstrate how estimating regret promote safety in real-world applications by abstention and data collection.

概率分类模型安全后悔值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。