arXiv:2512.24443cs.LGstat.ML2025-12

高维数据下用带置信度的正样本做稀疏分类,性能接近全监督方法。

Sparse classification with positive-confidence data in high dimensions

  • 设计基于L1、SCAD、MCP惩罚的稀疏框架,缓解高维下特征估计偏差。
  • 理论证明在受限强凸条件下达到近最优稀疏恢复率,误差可量化。
  • 适合缺乏负样本但特征远多于样本的高维分类任务,如生物医学分析。

当特征数量超过样本量时,高维学习常需稀疏正则化以实现有效预测与变量选择。尽管全监督场景下已有成熟方法,但在弱监督设置如正置信(Pconf)分类中仍研究不足。Pconf仅使用带置信度的正样本,无需负样本。现有Pconf方法在高维场景表现不佳。本文提出一种新的高维Pconf稀疏惩罚框架,采用凸(Lasso)与非凸(SCAD、MCP)惩罚,以缓解收缩偏差并提升特征恢复能力。理论上,我们建立了L1正则化Pconf估计器的估计与预测误差界,证明其在受限强凸条件下达到近最小极大最优稀疏恢复率。为求解复合目标函数,开发了高效的近端梯度算法。大量模拟实验表明,所提方法在预测性能与变量选择精度上均与全监督方法相当,有效弥合弱监督与高维统计之间的差距。

原文摘要 · Abstract (English)

High-dimensional learning problems, where the number of features exceeds the sample size, often require sparse regularization for effective prediction and variable selection. While established for fully supervised data, these techniques remain underexplored in weak-supervision settings such as Positive-Confidence (Pconf) classification. Pconf learning utilizes only positive samples equipped with confidence scores, thereby avoiding the need for negative data. However, existing Pconf methods are ill-suited for high-dimensional regimes. This paper proposes a novel sparse-penalization framework for high-dimensional Pconf classification. We introduce estimators using convex (Lasso) and non-convex (SCAD, MCP) penalties to address shrinkage bias and improve feature recovery. Theoretically, we establish estimation and prediction error bounds for the L1-regularized Pconf estimator, proving it achieves near minimax-optimal sparse recovery rates under Restricted Strong Convexity condition. To solve the resulting composite objective, we develop an efficient proximal gradient algorithm. Extensive simulations demonstrate that our proposed methods achieve predictive performance and variable selection accuracy comparable to fully supervised approaches, effectively bridging the gap between weak supervision and high-dimensional statistics.

高维学习弱监督稀疏分类Pconf

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。