arXiv:2511.18773cs.LGcs.CV2025-11AAAI被引 1

针对半监督学习中类别不平衡问题,提出采样控制新框架提升模型准确性。

Sampling Control for Imbalanced Calibration in Semi-Supervised Learning

  • 分离采样控制机制,区分数据分布与学习难度差异影响。
  • 在训练和推理阶段分别优化特征与权重不平衡,提升少数类识别效果。
  • 适用于各类分布不匹配场景,尤其适合长尾数据集的半监督学习。

类别不平衡仍是半监督学习(SSL)中的关键挑战,尤其当标注数据与未标注数据分布不一致时易引发分类偏差。现有方法虽通过估计未标注数据的类别分布来调整输出概率,但通常粗粒度地处理模型偏差,混淆了数据分布不均与不同类别学习难度差异的影响。为此,我们提出统一框架SC-SSL,通过解耦采样控制抑制模型偏差。训练阶段,识别理想条件下采样控制的关键变量,引入具备显式扩展能力的分类器,并自适应调节不同数据分布下的采样概率,缓解少数类的特征级不平衡。推理阶段,进一步分析线性分类器的权重不平衡,通过后处理采样控制与优化偏差向量直接校准输出逻辑值。在多个基准数据集与分布设置下进行的大量实验验证了SC-SSL的一致性与先进性能。

原文摘要 · Abstract (English)

Class imbalance remains a critical challenge in semi-supervised learning (SSL), especially when distributional mismatches between labeled and unlabeled data lead to biased classification. Although existing methods address this issue by adjusting logits based on the estimated class distribution of unlabeled data, they often handle model imbalance in a coarse-grained manner, conflating data imbalance with bias arising from varying class-specific learning difficulties. To address this issue, we propose a unified framework, SC-SSL, which suppresses model bias through decoupled sampling control. During training, we identify the key variables for sampling control under ideal conditions. By introducing a classifier with explicit expansion capability and adaptively adjusting sampling probabilities across different data distributions, SC-SSL mitigates feature-level imbalance for minority classes. In the inference phase, we further analyze the weight imbalance of the linear classifier and apply post-hoc sampling control with an optimization bias vector to directly calibrate the logits. Extensive experiments across various benchmark datasets and distribution settings validate the consistency and state-of-the-art performance of SC-SSL.

半监督学习类别不平衡采样控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。