arXiv:2608.20667cs.LGcs.AI2026-08中稿 · IEEE ICSSE 2026 fo…

提出C-Score评估半监督学习在开放世界污染下的隐性退化,超越准确率局限。

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

  • 从预测、特征和优化三方面诊断伪标签学习的内部行为
  • 在SVHN污染下CCI上升280%,准确率却仅下降3%
  • 适合关注模型鲁棒性与内部机制分析的研究者

基于伪标签的半监督学习因简洁性和可扩展性表现优异,但通常假设未标记数据与标记数据来自相同分布。实际部署中,未标记数据常来自开放环境,可能包含分布外(OOD)样本。这些样本可能获得高置信度预测并被错误纳入训练。此时,干净的分布内测试准确率可能保持稳定,但模型内部学习已退化。本文从诊断评估角度研究伪标签方法在开放世界未标记污染下的隐性崩溃现象,提出C-Score框架,在预测、特征表示和优化三个空间评估训练行为:包含用于未标记预测行为的PLE和CCI,用于偏离标记语义锚点的Sem-Drift,以及用于标记与未标记优化兼容性的Grad-Align。在CIFAR-10和CIFAR-100上使用多个外部数据源、不同污染比例及四种伪标签方法的实验表明,C-Score能揭示准确率无法检测的隐性退化:在SVHN污染下,CCI上升超280%,最佳准确率仍仅比无污染基线低3%;近分布外来源(如CIFAR-100、STL-10)导致最大14.9%准确率下降(FlexMatch,r=0.5)。结果表明,仅依赖准确率不足以评估开放世界下半监督学习的鲁棒性,需结合内部诊断信号进行更可靠的评估。

原文摘要 · Abstract (English)

Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and scalability. However, it is typically developed under a closed-world assumption that unlabeled data are drawn from the same distribution as labeled data. In practical deployment, unlabeled data are often collected from open environments and may contain OOD samples. Under such contamination, OOD samples may still receive high-confidence predictions and be incorporated into training as if they were valid target examples. This creates an important evaluation problem: clean in-distribution test accuracy may appear stable even when the internal learning dynamics of SSL have already deteriorated. To address this issue, we study hidden collapse in pseudo-label-based SSL under open-world unlabeled contamination from a diagnostic evaluation perspective. We present C-Score, a compact framework that evaluates training behavior in three complementary spaces: prediction, feature representation, and optimization. C-Score includes PLE and CCI for unlabeled prediction behavior, Sem-Drift for deviation from labeled semantic anchors, and Grad-Align for the compatibility between labeled and unlabeled optimization. Experiments on CIFAR-10 and CIFAR-100 with multiple OOD sources, varying contamination ratios, and four pseudo-label-based SSL algorithms show that C-Score metrics reveal hidden degradation that clean accuracy alone fails to detect: under SVHN contamination, CCI rises over 280% while best-accuracy remains within 3% of the uncontaminated baseline; near-OOD sources (CIFAR-100, STL-10) cause up to 14.9% accuracy collapse (FlexMatch, r=0.5). The results suggest that clean accuracy alone is insufficient for evaluating SSL robustness in open-world environments, and that internal diagnostic signals are necessary for more reliable robustness assessment under unlabeled contamination.

半监督学习鲁棒性评估开放世界诊断指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。