高维稀疏分类中,联合使用标签与无标签数据可显著提升模型准确率。
Semi-Supervised Sparse Gaussian Classification: Provable Benefits of Unlabeled Data
- 基于高维稀疏高斯分类,设计可证明有效的半监督学习方法
- 在特定参数范围内,仅用标签或无标签数据的算法均失效,而半监督方法有效
- 理论证明了无标签数据在特征选择与分类中的不可替代价值
半监督学习(SSL)的核心假设是:结合标签数据与无标签数据能显著提升模型准确性。尽管实证上已取得成功,但其理论理解仍不充分。本文研究高维稀疏高斯分类中的半监督学习问题。关键任务是特征选择——识别出少数能区分两类的关键变量。我们分析了在信息论和计算复杂度上的下界,假设低阶似然困难性猜想成立。主要贡献在于识别出一个参数区间(维度、稀疏度、标签与无标签样本数量),在此区间内,半监督学习在分类上具有确定优势。具体而言,在多项式时间内可构造出准确的半监督分类器,而任何仅依赖标签或仅依赖无标签数据的高效学习方法都将失败。本工作揭示了在高维场景下,联合利用标签与无标签数据在分类与特征选择中的可证明优势。实验模拟验证了理论分析。
原文摘要 · Abstract (English)
The premise of semi-supervised learning (SSL) is that combining labeled and unlabeled data yields significantly more accurate models. Despite empirical successes, the theoretical understanding of SSL is still far from complete. In this work, we study SSL for high dimensional sparse Gaussian classification. To construct an accurate classifier a key task is feature selection, detecting the few variables that separate the two classes. % For this SSL setting, we analyze information theoretic lower bounds for accurate feature selection as well as computational lower bounds, assuming the low-degree likelihood hardness conjecture. % Our key contribution is the identification of a regime in the problem parameters (dimension, sparsity, number of labeled and unlabeled samples) where SSL is guaranteed to be advantageous for classification. Specifically, there is a regime where it is possible to construct in polynomial time an accurate SSL classifier. However, % any computationally efficient supervised or unsupervised learning schemes, that separately use only the labeled or unlabeled data would fail. Our work highlights the provable benefits of combining labeled and unlabeled data for {classification and} feature selection in high dimensions. We present simulations that complement our theoretical analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。