提出可计算最优正则化强度的迭代方法,提升线性回归泛化性能。
Optimal ridge regularization revisited

- 基于生成参数设计迭代算法,求解固定数据下的最优正则化强度。
- 在不同样本量、噪声水平下,接近最优随机数据泛化误差。
- 计算成本低,仅需一次或两次岭回归开销,适合实际应用。
考虑在有限数据样本 $X$ 上进行 $L^2$ 正则化的线性(岭)回归,假设协方差有界,预测目标 $y$ 存在有限方差的各向同性加性噪声。本文提出一种迭代算法,可在固定 $X$ 设置下从生成参数数值计算最优正则化强度,并证明其在低噪声水平下收敛。对合成数据的实验表明,该方法结合基于样本的参数估计,在广泛样本规模、纵横比和噪声水平下实现接近最优的随机 $X$ 泛化性能,计算开销仅相当于在欠参数情形下执行一次预估岭回归,过参数情形下执行两次。
原文摘要 · Abstract (English)
We consider $L^2$-regularized linear (ridge) regression over a finite data sample $X$ with bounded covariance and linear prediction targets $y$ with additive isotropic noise of finite variance. We present an iterative procedure to compute the optimal regularization strength numerically from the generative parameters in the fixed-$X$ setting and prove its convergence at limited noise levels. Our experimental evaluation over synthetic data shows that the proposed procedure combined with sample-based parameter estimates attains near-optimal random-$X$ generalization across a wide range of sample sizes, aspect ratios, and noise levels, at an added computational cost equivalent to one preliminary ridge regression in the underparameterized regime and two in the overparameterized case.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。