通过时间一致性提升自训练在分布偏移下的性能,无需额外计算开销。
Improving self-training under distribution shifts via anchored confidence with theoretical guarantees
- 基于不确定性感知的时间集成与相对阈值法,增强伪标签可靠性。
- 在多种分布偏移场景下,性能提升8%至16%,且渐近正确。
- 方法鲁棒性强,对超参数不敏感,适合实际部署场景。
自训练在分布偏移下表现不佳,主要因预测置信度与实际准确率之间差距增大。传统方法如邻域或集成标签修正需大量计算。受早期学习正则化启发,本文提出一种基于时间一致性的原理性改进方法:构建不确定性感知的时间集成,并采用简单相对阈值策略;该集成可平滑噪声伪标签,促进选择性时间一致性。理论证明该时间集成渐近正确,标签平滑技术可缩小自训练的最优性差距。大量实验表明,本方法在多种分布偏移场景下均实现8%~16%的性能提升,且无计算开销增加。此外,方法具备更优校准性能和对超参数变化的强鲁棒性。
原文摘要 · Abstract (English)
Self-training often falls short under distribution shifts due to an increased discrepancy between prediction confidence and actual accuracy. This typically necessitates computationally demanding methods such as neighborhood or ensemble-based label corrections. Drawing inspiration from insights on early learning regularization, we develop a principled method to improve self-training under distribution shifts based on temporal consistency. Specifically, we build an uncertainty-aware temporal ensemble with a simple relative thresholding. Then, this ensemble smooths noisy pseudo labels to promote selective temporal consistency. We show that our temporal ensemble is asymptotically correct and our label smoothing technique can reduce the optimality gap of self-training. Our extensive experiments validate that our approach consistently improves self-training performances by 8% to 16% across diverse distribution shift scenarios without a computational overhead. Besides, our method exhibits attractive properties, such as improved calibration performance and robustness to different hyperparameter choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。