arXiv:2411.19553cs.LGstat.ML2024-11

分析高维高斯混合模型中半监督学习的性能边界与正则化作用。

Analysis of High-dimensional Gaussian Labeled-unlabeled Mixture Model via Message-passing Algorithm

  • 用消息传递算法研究高维二分类半监督学习的理论性质。
  • ℓ₂正则化使非贝叶斯方法接近最优,尤其在大量无标签数据时。
  • 对比贝叶斯最优与最大似然估计,揭示正则化关键作用。

半监督学习(SSL)利用少量标注数据和大量未标注数据提升模型性能,但其何时有效仍缺乏理论理解。已有研究通过高斯混合模型(GMM)建模分类问题,但分析多聚焦特定目标。本文针对高维二分类场景下的高维GMM,采用近似消息传递与状态演化方法,系统分析贝叶斯估计与ℓ₂正则化最大似然估计(RMLE)的表现。比较了全局相图、参数估计误差与标签预测误差。结果表明,在合适正则化下,RMLE在参数估计与标签预测误差上均能接近最优,尤其当无标签数据量大时表现优异。这证明ℓ₂正则化在半监督学习中对估计与预测均有显著提升作用。

原文摘要 · Abstract (English)

Semi-supervised learning (SSL) is a machine learning methodology that leverages unlabeled data in conjunction with a limited amount of labeled data. Although SSL has been applied in various applications and its effectiveness has been empirically demonstrated, it is still not fully understood when and why SSL performs well. Some existing theoretical studies have attempted to address this issue by modeling classification problems using the so-called Gaussian Mixture Model (GMM). These studies provide notable and insightful interpretations. However, their analyses are focused on specific purposes, and a thorough investigation of the properties of GMM in the context of SSL has been lacking. In this paper, we conduct such a detailed analysis of the properties of the high-dimensional GMM for binary classification in the SSL setting. To this end, we employ the approximate message passing and state evolution methods, which are widely used in high-dimensional settings and originate from statistical mechanics. We deal with two estimation approaches: the Bayesian one and the $\ell_2$-regularized maximum likelihood estimation (RMLE). We conduct a comprehensive comparison between these two approaches, examining aspects such as the global phase diagram, estimation error for the parameters, and prediction error for the labels. A specific comparison is made between the Bayes-optimal (BO) estimator and RMLE, as the BO setting provides optimal estimation performance and is ideal as a benchmark. Our analysis shows that with appropriate regularizations, RMLE can achieve near-optimal performance in terms of both the estimation error and prediction error, especially when there is a large amount of unlabeled data. These results demonstrate that the $\ell_2$ regularization term plays an effective role in estimation and prediction in SSL approaches.

半监督学习高斯混合模型正则化消息传递

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。