在标签稀缺时,用代理标签模型评估目标域表现,解决分布偏移问题。
Inferential Evaluation of Surrogate-Derived Models under Covariate Shift

- 利用双源密度比迁移信息,跨样本融合标签数据
- 在真实数据集上验证了模型在阈值附近的准确评估能力
- 适合医疗、金融等高风险场景的AI模型可靠性检验
在迁移学习中,基于丰富代理标签训练的模型可能部署于黄金标准标签不可得的目标群体。当源与目标数据分布不同时,评估其在目标域的表现极具挑战。本文研究三样本设置:少量黄金标签源、较大规模代理标签源和无标签目标。在条件可转移性假设下,通过交叉拟合估计器,利用源特定密度比将两标注源的信息迁移到目标。结合结果回归增强与核校正方法,估计模型在阈值附近的性能,同时考虑三个样本的不确定性。理论证明了真阳性率(TPR)和假阳性率(FPR)的渐近线性推断,受试者工作特征曲线(ROC)的一致性与逐点推断,以及曲线下面积(AUC)的渐近正态推断。模拟实验评估偏差、置信区间覆盖率及对带宽和样本量比例的敏感性。在Chatbot Arena的回顾性时间验证与半合成的ACS-Income研究中展示了实际应用有效性。
原文摘要 · Abstract (English)
In transfer-learning settings, a model derived from abundant surrogate labels may be deployed in a target population where gold-standard outcomes are unobserved. Evaluating its target performance is essential for determining whether decisions based on the model remain reliable, yet it is difficult when gold labels are scarce, and covariate distributions differ across data sources. We study a three-sample setting with a small gold-labeled source, a larger surrogate-labeled source, and an unlabeled target. Under conditional transportability, we evaluate the surrogate-derived model against the latent gold-standard outcome in the target population. We propose cross-fitted estimators that transport information from the two labeled sources through source-specific density ratios. We also combine outcome-regression augmentation with a kernel correction for estimating the model near a threshold, accounting for uncertainty from all three samples. We establish asymptotically linear inference for TPR and FPR, consistency and pointwise inference for the ROC curve, and asymptotically normal inference for AUC. Simulations assess bias, coverage, and sensitivity to bandwidth and relative sample sizes. A retrospective temporal validation on Chatbot Arena and a semi-synthetic ACS-Income study provide validation in real-world AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。