提出可预测置信度阈值偏差的诊断方法,判断伪标签回归是否可信。
When to Trust Confidence Thresholding: Calibration Diagnostics for Pseudo-Labelled Regression
- 基于残差得分方差构建校准感知诊断工具
- 推导出阈值导致的衰减偏差闭式表达式
- 提供可计算的决策规则,适合实证研究者使用
已训练分类器的校准概率输出被广泛用于下游回归估计,如潜在群体的效应、患病率或不平等性,这些群体仅在小规模标注子集上可观测。标准做法是将校准得分在置信度阈值处截断,将其硬标签视为真实标签。本文基于近期关于底层矩方程的识别结果,构建了一套校准敏感的伪标签管道诊断框架。我们推导出置信度阈值对下游回归系数造成的衰减偏差的闭式表达式,并证明该偏差可在任何推断前,通过在未标注集上偏移下游控制变量 $X$ 后的残差得分方差 $V^{*} = \mathbb{E}[\operatorname{Var}(p\mid X)]$ 预测。进一步,在受限校准漂移下获得紧致敏感性边界,并识别出 $V^{*}=0$ 的临界点——当且仅当 $p$ 是 $X$ 的确定性函数时成立,这揭示了分类器特征 $W$ 与下游控制变量 $X\subsetneq W$ 之间的结构分离。五组受控模拟及一个 UCI Adult 数据集示例验证了预测效果。本研究贡献在于提供了一个可操作的 $(V^{*}, \kappa)$ 决策规则,供从业者从任意分类器输出中计算,以判断置信度阈值是否安全。
原文摘要 · Abstract (English)
Calibrated probability outputs of trained classifiers are increasingly used as inputs to downstream regression estimands such as effects, prevalences, or disparities for a latent group observed only on a small labelled subset. A standard practice is to threshold the calibrated score at a confidence cutoff and treat the hard label as the truth. Building on a recent identification result for the underlying moment equation, we develop a calibration-aware diagnostic apparatus for pseudo-labelling pipelines. We derive a closed-form expression for the attenuation bias that confidence thresholding induces in the downstream regression coefficient, and show that the bias can be predicted, before any inference is run, from the residual score variance $V^{*}=\mathbb{E}[\operatorname{Var}(p\mid X)]$ on the unlabelled set after partialling out the downstream controls $X$. We further obtain a sharp sensitivity bound under bounded calibration drift, and identify the boundary $V^{*}=0$, which holds iff $p$ is a deterministic function of $X$; this motivates a structural separation between classifier features $W$ and downstream controls $X\subsetneq W$. Five controlled simulations and a UCI Adult illustration trace the predictions. The contribution is operational: a $(V^{*}, κ)$ decision rule that practitioners can compute from any classifier output to decide whether confidence thresholding is safe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。