提出新型稀疏回归模型,提升高维数据变量选择与预测精度。
Iterative Reweighted Framework Based Algorithms for Sparse Linear Regression with Generalized Elastic Net Penalty
- 用ℓ_q-范数替代L1,结合ℓ_r-范数增强噪声鲁棒性。
- 两种算法均在模拟与真实数据上表现更优,PMM-SSN效率更高。
- 适合处理高维稀疏回归问题的研究者与数据科学家。
弹性网惩罚常用于高维统计中的参数回归与变量选择,尤其当预测变量远多于观测数时表现更优。然而,实证表明ℓ_q-范数(0 < q < 1)相比ℓ_1-范数在多种场景下具有更强的鲁棒性与更好回归性能。本文提出一种广义弹性网模型:在损失函数中引入ℓ_r-范数(r ≥ 1)以适应不同噪声类型,并用ℓ_q-范数(0 < q < 1)替换原弹性网中的ℓ_1-范数。理论上,建立了广义一阶驻点非零项的可计算下界。为实现该模型,开发了基于局部Lipschitz连续ε-逼近ℓ_q-范数的两类高效算法:第一类采用交替方向乘子法(ADMM),第二类采用近端极大极小化方法(PMM),其中子问题通过半光滑牛顿法(SNN)求解。大量模拟与真实数据实验表明,两类算法性能均优于传统方法;值得注意的是,尽管ADMM实现更简单,但PMM-SNN更具效率。
原文摘要 · Abstract (English)
The elastic net penalty is frequently employed in high-dimensional statistics for parameter regression and variable selection. It is particularly beneficial compared to lasso when the number of predictors greatly surpasses the number of observations. However, empirical evidence has shown that the $\ell_q$-norm penalty (where $0 < q < 1$) often provides better regression compared to the $\ell_1$-norm penalty, demonstrating enhanced robustness in various scenarios. In this paper, we explore a generalized elastic net model that employs a $\ell_r$-norm (where $r \geq 1$) in loss function to accommodate various types of noise, and employs a $\ell_q$-norm (where $0 < q < 1$) to replace the $\ell_1$-norm in elastic net penalty. Theoretically, we establish the computable lower bounds for the nonzero entries of the generalized first-order stationary points of the proposed generalized elastic net model. For implementation, we develop two efficient algorithms based on the locally Lipschitz continuous $ε$-approximation to $\ell_q$-norm. The first algorithm employs an alternating direction method of multipliers (ADMM), while the second utilizes a proximal majorization-minimization method (PMM), where the subproblems are addressed using the semismooth Newton method (SNN). We also perform extensive numerical experiments with both simulated and real data, showing that both algorithms demonstrate superior performance. Notably, the PMM-SSN is efficient than ADMM, even though the latter provides a simpler implementation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。