arXiv:2409.11653cs.LGcs.CV2024-09NeurIPS被引 3

通过选代表性多样样本,提升低预算下的半监督学习效果。

Enhancing Semi-Supervised Learning via Representative and Diverse Sample Selection

  • 用改进的Frank-Wolfe算法最小化α-MMD,选样本
  • 在极低标注预算下显著提升模型泛化能力
  • 适合标注资源紧缺的半监督学习场景

半监督学习(SSL)已成为众多深度学习任务的首选范式,能减少人工标注成本。以往研究主要关注如何有效利用已标注和未标注数据以提升性能,但我们发现,在极端低预算设置下,样本选择方式对性能影响显著,而这一问题长期被忽视。为此,本文提出代表性和多样性样本选择方法(RDSS),采用改进的Frank-Wolfe算法最小化一种新准则——α-最大均值差异(α-MMD),从无标签数据中选出具有代表性和多样性的子集进行标注。实验表明,最小化α-MMD能增强低预算学习的泛化能力。在多个主流SSL框架上,RDSS均持续提升性能,并超越当前最先进的主动学习(AL)与半监督主动学习(SSAL)样本选择方法,即使在严格标注预算限制下依然表现优异。

原文摘要 · Abstract (English)

Semi-Supervised Learning (SSL) has become a preferred paradigm in many deep learning tasks, which reduces the need for human labor. Previous studies primarily focus on effectively utilising the labelled and unlabeled data to improve performance. However, we observe that how to select samples for labelling also significantly impacts performance, particularly under extremely low-budget settings. The sample selection task in SSL has been under-explored for a long time. To fill in this gap, we propose a Representative and Diverse Sample Selection approach (RDSS). By adopting a modified Frank-Wolfe algorithm to minimise a novel criterion $α$-Maximum Mean Discrepancy ($α$-MMD), RDSS samples a representative and diverse subset for annotation from the unlabeled data. We demonstrate that minimizing $α$-MMD enhances the generalization ability of low-budget learning. Experimental results show that RDSS consistently improves the performance of several popular SSL frameworks and outperforms the state-of-the-art sample selection approaches used in Active Learning (AL) and Semi-Supervised Active Learning (SSAL), even with constrained annotation budgets.

半监督学习样本选择低预算学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。