新方法无需假设数据分布,用样本关系提升半监督学习效果。
Beyond Distribution Estimation: Simplex Anchored Structural Inference Towards Universal Semi-Supervised Learning

- 通过高阶样本关系建模实现无分布假设的结构推断
- 在五个基准上平均提升8.52%准确率,显著优于现有方法
- 适合标签稀缺且数据分布未知的现实场景
半监督学习在真实场景中面临标注数据稀缺、无标注数据分布未知且任意的问题。本文将此关键但未被充分研究的范式形式化为通用半监督学习(UniSSL)。现有方法通常依赖伪标签,但往往假设无标注数据服从均匀分布或需要足够多标注数据来估计分布,导致大量错误伪标签,引发表示混淆。我们发现表示所捕捉的样本间关系比伪标签更可靠,因此转向表示层面的结构推断以避开分布估计。为此提出SAGE方法,通过捕获高阶样本依赖建立结构共识以指导表示学习;同时采用满足单纯形等角紧框架的向量作为坐标系,促进类间表示分离;最后引入与分布无关的度量加权策略筛选可靠伪标签,并设计辅助分支隔离潜在错误伪标签。在五个标准基准上的评估表明,SAGE始终优于当前最优方法,平均准确率提升8.52%。
原文摘要 · Abstract (English)
Semi-supervised learning faces significant challenges in realistic scenarios where labeled data is scarce and unlabeled data follows unknown, arbitrary distributions. We formalize this critical yet under-explored paradigm as Universal Semi-supervised Learning (UniSSL). Existing methods typically leverage unlabeled data via pseudo-labeling. However, they often rely on the idealized assumption of a uniform unlabeled data distribution or require sufficient labeled data to estimate it. In the UniSSL setting, such dependencies lead to numerous erroneous pseudo-labels, thereby triggering representation confusion. Fortunately, we observe that inter-sample relations captured by representations are more reliable than pseudo-labels. Leveraging this insight, we shift our focus to representation-level structural inference to bypass distribution estimation. Accordingly, we propose Simplex Anchored Graph-state Equipartition (SAGE), which captures high-order inter-sample dependencies to establish structural consensus for guiding representation learning. Meanwhile, to mitigate representation confusion, we employ vectors that satisfy a simplex equiangular tight frame to serve as a coordinate frame for guiding inter-class representation separation. Finally, we introduce a weighting strategy based on distribution-agnostic metrics to prioritize reliable pseudo-labels and an auxiliary branch to isolate potentially erroneous pseudo-labels. Evaluations on five standard benchmarks show that SAGE consistently outperforms state-of-the-art methods, with an average accuracy gain of $\textbf{8.52%}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。