提出可直接计算的正则化参数,避免交叉验证且提升稀疏精度矩阵估计效率。
The Regularization Parameter: Sparse Precision Matrix Estimation

- 基于一阶最优性条件的采样分布,推导出闭式矩阵形式的正则化参数。
- 在合成与真实数据上准确率媲美交叉验证,支持恢复性能更优,速度提升数个数量级。
- 适合高维低样本场景,如基因芯片和脑成像分析,无需调参耗时。
稀疏精度矩阵估计为高维、小样本数据中的条件依赖建模提供了可解释且计算高效的框架。一个持续的挑战是如何恰当选择控制估计器稀疏性的正则化参数,以平衡欠拟合与过拟合。本文提出一种从ℓ₁-正则化高斯最大似然估计的一阶最优性条件采样分布中导出的闭式、矩阵型正则化参数。通过规定每个非零项在重抽样下满足最优性条件的概率,消除了对交叉验证的需求。所提出的正则化参数展现出渐近标度性质,在标准条件下保证估计器的一致性和稀疏一致性。在合成高斯与非高斯数据,以及真实世界的基因微阵列和神经影像应用中,该方法达到与交叉验证相当的估计精度,提供更优的支持恢复效果,并将运行时间减少数个数量级。
原文摘要 · Abstract (English)
Sparse precision matrix estimation provides an interpretable and computationally efficient framework for modeling conditional dependencies in high-dimensional, low-sample-size data. A recurring challenge is appropriately selecting the regularization parameter that controls estimator sparsity and strikes a balance between underfitting and overfitting. We propose a closed-form, matrix-valued regularization parameter derived from the sampling distribution of the first-order optimality conditions of the $\ell_1$-regularized Gaussian maximum-likelihood estimator. By prescribing the probability that each nonzero entry of the estimator satisfies its optimality condition under resampling, we eliminate the need for cross-validation. The resulting regularization parameter is shown to attain asymptotic scaling properties that, under standard conditions, provide consistency and sparsistency of the estimator. On synthetic Gaussian and non-Gaussian datasets, as well as real-world gene microarray and neuroimaging applications, the proposed approach achieves estimation accuracy comparable to cross-validation, delivers superior support recovery, and reduces runtime by several orders of magnitude.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。