用可随时验证的方法,检验标签偏移修正是否合理。
Anytime-Valid Confirmation of Label-Shift Corrections
- 基于似然比构建条件 e 值,形成可随时检验的统计过程。
- 实测在小批量场景下,先验修正比重新估计更可靠。
- 适合需要持续监控模型稳定性的科研部署场景。
在小批量科学部署中,目标域标注样本过少,难以可靠估计分布偏移,即使有未标注目标输入可用。本文研究互补场景:从业者已有基于领域知识的预设标签偏移修正,需判断新到来的标注数据是否支持该修正。我们证明,标签修正预测与源预测之间的逐样本似然比是条件 e 值,其累积乘积构成非负鞅,结合 Ville 不等式可得任意时间有效的确认规则。对数鞅等于源预测与修正预测之间累积负对数概率密度(NLPD)差距,使常规模型监控转化为正式的序列检验。拒绝原假设意味着新数据支持该修正,但不精确反映偏移程度。对于高斯过程(GP)源及高斯标签偏移比的情形,存在闭式解。模拟实验验证了第一类错误控制、有限样本功效、误校准敏感性,以及小批量下先验修正优于基于标签重估计的优势。
原文摘要 · Abstract (English)
In small-batch scientific deployments, labeled target outcomes may be too scarce for reliable shift estimation even when unlabeled target inputs are available. We address the complementary setting where the practitioner has a pre-specified label-shift correction from domain knowledge and asks whether incoming labeled outcomes support it. We show that the per-observation likelihood ratio between a label-shift-corrected predictive and the source predictive is a conditional e-value, so its running product is a nonnegative martingale and Ville's inequality yields an anytime-valid confirmation rule. The log martingale equals the cumulative negative log-predictive density (NLPD) gap between the source and the corrected predictive, converting routine model monitoring into a formal sequential test. Rejection means the incoming data support the posited correction relative to the source predictive, but it is not a precise estimate of the degree of shift. Closed forms are available for GP sources with Gaussian label-shift ratios. GP regression simulations validate Type I control, finite-sample power, miscalibration sensitivity, and the small-batch advantage of a reliable prior over label-based re-estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。