用自适应阈值提升少标签场景下的域泛化能力
CAT: Class Aware Adaptive Thresholding for Semi-Supervised Domain Generalization
- 动态调整伪标签阈值,增强类别多样性
- 通过噪声伪标签修正,显著提高标签可靠性
- 适合标注数据稀缺的现实场景应用
域泛化(DG)旨在将多个源域的知识迁移至未见的目标域,即使存在域偏移。传统方法通常依赖大量多样化的标注源数据以学习可泛化的表征,但高质量标注数据获取成本高,限制了实际应用。为此,本文研究更贴近实际且更具挑战性的问题:在标签高效范式下的半监督域泛化(SSDG)。提出新方法CAT,利用有限标注数据结合半监督学习,在域偏移下实现优异的泛化性能。相比先前方法对固定阈值的依赖及对噪声伪标签的敏感性,CAT融合自适应阈值与噪声标签优化技术,通过灵活阈值生成更具类多样性的高质量伪标签,并进一步修正噪声标签以提升其可信度。在多个基准数据集上的广泛实验表明,该方法在域偏移下表现卓越,验证了其在鲁棒泛化方面的有效性。
原文摘要 · Abstract (English)
Domain Generalization (DG) seeks to transfer knowledge from multiple source domains to unseen target domains, even in the presence of domain shifts. Achieving effective generalization typically requires a large and diverse set of labeled source data to learn robust representations that can generalize to new, unseen domains. However, obtaining such high-quality labeled data is often costly and labor-intensive, limiting the practical applicability of DG. To address this, we investigate a more practical and challenging problem: semi-supervised domain generalization (SSDG) under a label-efficient paradigm. In this paper, we propose a novel method, CAT, which leverages semi-supervised learning with limited labeled data to achieve competitive generalization performance under domain shifts. Our method addresses key limitations of previous approaches, such as reliance on fixed thresholds and sensitivity to noisy pseudo-labels. CAT combines adaptive thresholding with noisy label refinement techniques, creating a straightforward yet highly effective solution for SSDG tasks. Specifically, our approach uses flexible thresholding to generate high-quality pseudo-labels with higher class diversity while refining noisy pseudo-labels to improve their reliability. Extensive experiments across multiple benchmark datasets demonstrate the superior performance of our method, highlighting its effectiveness in achieving robust generalization under domain shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。