通过挖掘难例提升不平衡半监督学习性能
SeMi: When Imbalanced Semi-Supervised Learning Meets Mining Hard Examples
- 基于置信度差异识别难例,增强未标注数据利用
- 使用带衰减的平衡记忆库提升伪标签可靠性
- 简单有效,尤其在类别反转场景下提升超54%
半监督学习(SSL)可利用大量未标注数据提升模型性能。然而,现实场景中的类别不平衡问题严重制约了SSL表现。现有类别不平衡半监督学习(CISSL)方法主要聚焦于数据重平衡,却忽视了难例挖掘对性能的潜力,导致即便使用复杂算法也难以充分释放未标注数据价值。为此,我们提出一种通过挖掘难例增强不平衡半监督学习性能的方法(SeMi)。该方法通过分析难例与易例之间logits的熵差异来识别难例,从而提升未标注数据利用率,更好解决CISSL中的不平衡问题。此外,我们构建了一个带有置信度衰减机制的类别平衡记忆库,用于存储高置信度嵌入,以增强伪标签可靠性。尽管方法简单,但效果显著,且能无缝集成到现有框架中。我们在标准CISSL基准上进行大量实验,结果表明,所提SeMi在多个基准上优于现有最先进方法,尤其在类别反转场景下,最佳结果较基线方法提升约54.8%。
原文摘要 · Abstract (English)
Semi-Supervised Learning (SSL) can leverage abundant unlabeled data to boost model performance. However, the class-imbalanced data distribution in real-world scenarios poses great challenges to SSL, resulting in performance degradation. Existing class-imbalanced semi-supervised learning (CISSL) methods mainly focus on rebalancing datasets but ignore the potential of using hard examples to enhance performance, making it difficult to fully harness the power of unlabeled data even with sophisticated algorithms. To address this issue, we propose a method that enhances the performance of Imbalanced Semi-Supervised Learning by Mining Hard Examples (SeMi). This method distinguishes the entropy differences among logits of hard and easy examples, thereby identifying hard examples and increasing the utility of unlabeled data, better addressing the imbalance problem in CISSL. In addition, we maintain a class-balanced memory bank with confidence decay for storing high-confidence embeddings to enhance the pseudo-labels' reliability. Although our method is simple, it is effective and seamlessly integrates with existing approaches. We perform comprehensive experiments on standard CISSL benchmarks and experimentally demonstrate that our proposed SeMi outperforms existing state-of-the-art methods on multiple benchmarks, especially in reversed scenarios, where our best result shows approximately a 54.8\% improvement over the baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。