提出可扩展到多分类的无分布假设半监督学习新框架
Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite

- 用线性组合构建无偏风险估计器,泛化了原有方法
- 理论证明在非对称损失下方差更低,提升学习性能
- 适用于二分类和多分类任务,实测效果优于现有方法
传统半监督学习方法依赖分布假设,一旦违反性能下降。虽有PNU学习提供无分布替代方案,但仅限二分类且方差最优性不明。本文提出广义框架,通过线性组合构建无偏风险估计器,涵盖PNU并扩展至多分类。推导出最小可实现方差,证明该估计器在非对称损失场景下方差低于PNU。进一步建立泛化界,直接关联方差降低与性能提升。基于此,设计两种实用半监督方法,在二分类和多分类基准上表现匹配或超越现有方法。
原文摘要 · Abstract (English)
Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated. While PNU learning, a risk rewriting method, offers a distribution-free alternative, it is restricted to binary classification and its variance optimality remains unclear. In this paper, we propose a generalized framework that constructs unbiased risk estimators using linear combinations of component risks, subsuming PNU learning and extending to multiclass classification. We derive the minimum achievable variance, demonstrating our estimator can attain lower variance than PNU in asymmetric loss scenarios. Furthermore, we establish a generalization bound directly linking this variance reduction to improved learning performance. Based on these theoretical insights, we introduce two practical SSL methods that empirically match or outperform existing approaches on binary and multiclass benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。