解决跨域半监督学习中标签分布差异问题
A Unified Framework for Heterogeneous Semi-supervised Learning
- 统一框架直接从异质数据中学习细粒度分类器
- 在多个数据集上超越现有半监督与无监督域适应方法
- 适合处理标签和特征分布不同的跨域学习场景
本文提出异质半监督学习(HSSL)新范式,将半监督学习(SSL)与无监督域适应(UDA)相结合,针对来自不同领域但共享语义类别的异质训练数据进行建模。该任务要求模型能区分来自有标签和无标签领域的测试样本,而这两个领域存在显著的标签分布和类别特征分布差异。为应对这一挑战,我们提出统一框架Uni-HSSL,通过自适应地处理域间异质性,同时利用无标签数据和跨域语义类别关系实现知识迁移与适应。实验表明,该方法在多个基准数据集上优于当前最先进的半监督学习和无监督域适应方法。
原文摘要 · Abstract (English)
In this work, we introduce a novel problem setup termed as Heterogeneous Semi-Supervised Learning (HSSL), which presents unique challenges by bridging the semi-supervised learning (SSL) task and the unsupervised domain adaptation (UDA) task, and expanding standard semi-supervised learning to cope with heterogeneous training data. At its core, HSSL aims to learn a prediction model using a combination of labeled and unlabeled training data drawn separately from heterogeneous domains that share a common set of semantic categories; this model is intended to differentiate the semantic categories of test instances sampled from both the labeled and unlabeled domains. In particular, the labeled and unlabeled domains have dissimilar label distributions and class feature distributions. This heterogeneity, coupled with the assorted sources of the test data, introduces significant challenges to standard SSL and UDA methods. Therefore, we propose a novel method, Unified Framework for Heterogeneous Semi-supervised Learning (Uni-HSSL), to address HSSL by directly learning a fine-grained classifier from the heterogeneous data, which adaptively handles the inter-domain heterogeneity while leveraging both the unlabeled data and the inter-domain semantic class relationships for cross-domain knowledge transfer and adaptation. We conduct comprehensive experiments and the experimental results validate the efficacy and superior performance of the proposed Uni-HSSL over state-of-the-art semi-supervised learning and unsupervised domain adaptation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。