arXiv:2410.19950stat.MLcs.LG2024-10

研究高维高斯混合分类的泛化误差与变量选择,提出新估计方法。

Statistical Inference in Classification of High-dimensional Gaussian Mixture

  • 用统计物理的副本法分析高维极限下正则化分类器行为。
  • 提出去偏估计器,可准确识别重要变量,错误率随维度上升仍可控。
  • 适用于高维生物数据、金融建模等需特征筛选的场景。

研究高维两分类高斯混合模型在一般协方差结构下的分类问题。通过统计物理中的副本法,分析当样本量n与维度p同时趋于无穷、其比值α = n/p保持固定时,一类广义正则化凸分类器的渐近性质,重点关注泛化误差与变量选择能力。基于分类器的分布极限,构造去偏估计器,通过合适的假设检验实现变量选择。以L1正则化逻辑回归为例,大量数值实验验证了理论分析结果在有限系统中依然成立。还探讨了协方差结构对去偏估计器性能的影响。

原文摘要 · Abstract (English)

We consider the classification problem of a high-dimensional mixture of two Gaussians with general covariance matrices. Using the replica method from statistical physics, we investigate the asymptotic behavior of a general class of regularized convex classifiers in the high-dimensional limit, where both the sample size $n$ and the dimension $p$ approach infinity while their ratio $α=n/p$ remains fixed. Our focus is on the generalization error and variable selection properties of the estimators. Specifically, based on the distributional limit of the classifier, we construct a de-biased estimator to perform variable selection through an appropriate hypothesis testing procedure. Using $L_1$-regularized logistic regression as an example, we conducted extensive computational experiments to confirm that our analytical findings are consistent with numerical simulations in finite-sized systems. We also explore the influence of the covariance structure on the performance of the de-biased estimator.

高维统计分类器变量选择去偏估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。