通过集成学习提升半监督域适应性能,无需重新训练
Semi-Supervised Transfer Boosting (SS-TrBoosting)
- 用提升法生成监督与半监督两类基学习器并集成
- 在多个基准上显著优于现有UDA/SSDA方法
- 支持无源数据场景,兼顾效率与隐私保护
半监督域适应(SSDA)旨在利用少量目标域标注数据、大量未标注目标数据和丰富的源域辅助数据,训练高性能模型。以往工作多聚焦于跨域可迁移表征学习,但难以找到源域与目标域条件分布一致的特征空间。同时,缺乏灵活有效的策略将现有无监督域适应(UDA)方法扩展至SSDA场景。为此,本文提出一种新的微调框架——半监督迁移增强(SS-TrBoosting)。给定一个已训练好的基于深度学习的UDA或SSDA模型,将其作为初始模型,通过提升法生成额外的基学习器,并将所有学习器组成集成模型。其中一半基学习器来自监督域适应,另一半来自半监督学习。为进一步提升数据传输效率并加强数据隐私保护,还提出了源数据生成方法,使该框架可扩展至半监督无源域适应(SS-SFDA)场景。大量实验表明,SS-TrBoosting可有效应用于多种现有UDA、SSDA和SFDA方法,持续提升其性能。
原文摘要 · Abstract (English)
Semi-supervised domain adaptation (SSDA) aims at training a high-performance model for a target domain using few labeled target data, many unlabeled target data, and plenty of auxiliary data from a source domain. Previous works in SSDA mainly focused on learning transferable representations across domains. However, it is difficult to find a feature space where the source and target domains share the same conditional probability distribution. Additionally, there is no flexible and effective strategy extending existing unsupervised domain adaptation (UDA) approaches to SSDA settings. In order to solve the above two challenges, we propose a novel fine-tuning framework, semi-supervised transfer boosting (SS-TrBoosting). Given a well-trained deep learning-based UDA or SSDA model, we use it as the initial model, generate additional base learners by boosting, and then use all of them as an ensemble. More specifically, half of the base learners are generated by supervised domain adaptation, and half by semi-supervised learning. Furthermore, for more efficient data transmission and better data privacy protection, we propose a source data generation approach to extend SS-TrBoosting to semi-supervised source-free domain adaptation (SS-SFDA). Extensive experiments showed that SS-TrBoosting can be applied to a variety of existing UDA, SSDA and SFDA approaches to further improve their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。