arXiv:2412.10741cs.LGcs.CV2024-12AAAI被引 9

改进Mixup在半监督学习中的使用,提升低置信度样本利用率

RegMixMatch: Optimizing Mixup Utilization in Semi-Supervised Learning

  • 采用混合样本与干净样本联合训练,提升人工标签纯净度
  • 对低置信度样本引入前两预测类信息,减少偏差并增强利用
  • 在多个基准上达到当前最优效果,适合高噪声数据场景

一致性正则化和伪标签已显著推动半监督学习(SSL)发展。以往工作有效将Mixup用于一致性正则化,但我们的研究发现,该做法可能因降低人工标签纯度而损害SSL性能。此外,多数伪标签方法通过阈值筛选低置信度数据以缓解确认偏差,但这限制了未标记样本的利用。为此,我们提出RegMixMatch框架,优化混合样本在高/低置信度数据上的应用。首先,引入半监督RegMixup,通过混合样本与干净样本共同训练,有效缓解人工标签纯度下降问题;其次,设计类别感知的Mixup技术,将前两预测类别信息融入低置信度样本及其人工标签,降低其带来的确认偏差并提升利用效率。实验表明,RegMixMatch在多个主流SSL基准上均达到当前最优性能。

原文摘要 · Abstract (English)

Consistency regularization and pseudo-labeling have significantly advanced semi-supervised learning (SSL). Prior works have effectively employed Mixup for consistency regularization in SSL. However, our findings indicate that applying Mixup for consistency regularization may degrade SSL performance by compromising the purity of artificial labels. Moreover, most pseudo-labeling based methods utilize thresholding strategy to exclude low-confidence data, aiming to mitigate confirmation bias; however, this approach limits the utility of unlabeled samples. To address these challenges, we propose RegMixMatch, a novel framework that optimizes the use of Mixup with both high- and low-confidence samples in SSL. First, we introduce semi-supervised RegMixup, which effectively addresses reduced artificial labels purity by using both mixed samples and clean samples for training. Second, we develop a class-aware Mixup technique that integrates information from the top-2 predicted classes into low-confidence samples and their artificial labels, reducing the confirmation bias associated with these samples and enhancing their effective utilization. Experimental results demonstrate that RegMixMatch achieves state-of-the-art performance across various SSL benchmarks.

半监督学习Mixup伪标签标签纯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。