利用无标签数据校准,安全利用不稳定的特征提升模型泛化能力
Adapting to Shifting Correlations with Unlabeled Data Calibration
- 基于新场景无标签数据推断目标与混淆因子的交互关系
- 在多个真实和合成数据集上超越现有基线方法
- 适用于存在多重混淆因素且需应对分布偏移的场景
不同站点之间的分布偏移会严重降低模型性能,因为模型容易依赖不稳定的关联。现有方法通常寻找跨站点稳定的特征并丢弃不稳定特征,但这些不稳定特征可能包含互补信息,若合理利用可提升准确率。近期方法尝试在新站点适应不稳定特征以提高精度,但常有不切实际假设或难以扩展至多混淆因子场景。本文提出广义流行度调整(GPA),一种灵活的方法,通过调整模型预测以适应目标与混淆因子间相关性的变化,从而安全地利用不稳定特征。GPA 利用新站点的无标签样本推断目标与混淆因子的交互关系。我们在多个真实和合成数据集上评估 GPA,结果表明其优于多种竞争基线方法。
原文摘要 · Abstract (English)
Distribution shifts between sites can seriously degrade model performance since models are prone to exploiting unstable correlations. Thus, many methods try to find features that are stable across sites and discard unstable features. However, unstable features might have complementary information that, if used appropriately, could increase accuracy. More recent methods try to adapt to unstable features at the new sites to achieve higher accuracy. However, they make unrealistic assumptions or fail to scale to multiple confounding features. We propose Generalized Prevalence Adjustment (GPA for short), a flexible method that adjusts model predictions to the shifting correlations between prediction target and confounders to safely exploit unstable features. GPA can infer the interaction between target and confounders in new sites using unlabeled samples from those sites. We evaluate GPA on several real and synthetic datasets, and show that it outperforms competitive baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。