提出新方法识别领域迁移中的私有类别,提升极端标签偏移下的分类准确率。
Locality-aware Private Class Identification for Domain Adaptation with Extreme Label Shift

- 基于最优传输的局部性设计得分函数,精准区分共享与私有类别样本。
- 在极端标签偏移下,新方法使目标域分类误差显著降低,实验验证有效。
- 适合处理标签分布差异大、存在私有类别的实际场景,如跨领域图像识别。
领域自适应旨在将标注源域的知识迁移到标签未标注的目标域,但二者分布不同。现实场景中,两域标签空间常存在包含关系,部分类别仅存在于一域,称为私有类别。现有方法假设私有类别差异足够大可视为异常点,但单个共享类内部方差可能远大于私有类与共享类间差异,挑战该假设。为此,本文基于最优传输(OT)的局部运输与度量特性,提出一种局域感知的私有类别识别方法,通过运输质量得分函数实现。理论证明该得分函数能有效区分共享与私有样本。在此基础上,构建可靠OT方法(ReOT),在严重标签偏移下最小化分类风险,并学习分离的簇结构,避免共享-私有样本对错配,确保知识可靠类内传输,缓解条件类别差异。进一步给出极端标签偏移下目标风险的泛化上界,可通过ReOT最小化。大量基准测试验证了ReOT的有效性。
原文摘要 · Abstract (English)
Domain adaptation aims to transfer knowledge from a labeled source domain to an unlabeled target domain with different distributions. In real-world scenarios, the label spaces of the two domains often have an inclusion relationship, where some classes exist only in one domain but not the other. These non-overlapping classes are referred to as private classes. Identifying private class samples and mitigating their adverse effects is critical in the literature. Existing methods rely on the assumption that shifts in private classes are large enough to be considered outliers. However, the variance within a single shared class can be significantly larger than the difference between a private class and another shared class, challenging this assumption. Consequently, private classes substantially increase the difficulty of cross-domain classification. To address these issues, based on local transportation and metric properties of optimal transport (OT), a locality-aware private class identification approach is proposed in the form of a score function on transport mass. The effectiveness of the proposed approach is theoretically proven, highlighting the score function's strong ability to distinguish between shared and private class samples. Building on this, we introduce a reliable OT-based method (ReOT) for domain adaptation under severe label shift. ReOT minimizes classification risk while learning the separated cluster structure between the identified shared classes and private classes, effectively avoiding mismatch between shared-private sample pairs, thus ensuring that important knowledge is reliably transported intra-class to mitigate class-conditional discrepancy. Furthermore, a generalization upper bound of the target risk is provided for extreme label shift scenarios, which can be minimized by ReOT. Extensive experiments on benchmarks validate the effectiveness of ReOT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。