解决高维回归中主成分回归的截断偏差问题
Calibrated Principal Component Regression
- 在主成分空间学先验,再通过中心化Tikhonov正则化校准原空间模型
- 理论证明在随机矩阵框架下,预测风险低于传统PCR
- 适合处理高维数据且信号分布不集中于前几个主成分的场景
我们提出一种广义线性模型中的统计推断新方法。在过参数化情形下,主成分回归(PCR)通过将高维数据投影到低维主成分子空间来降低方差,但当真实回归向量在保留的主成分之外有显著分量时,会引入截断偏差。为缓解此问题,我们提出校准主成分回归(CPCR),先在主成分子空间学习低方差先验,再通过中心化Tikhonov步骤在原始特征空间进行校准。CPCR结合交叉拟合,通过软化PCR的硬截断来控制截断偏差。理论上,我们在随机矩阵框架下计算了样本外风险,结果表明当回归信号在低方差方向有非可忽略分量时,CPCR优于标准PCR。实证上,CPCR在多个过参数化问题中均稳定提升预测性能,凸显其在现代过参数化设定下的稳定性和灵活性。
原文摘要 · Abstract (English)
We propose a new method for statistical inference in generalized linear models. In the overparameterized regime, Principal Component Regression (PCR) reduces variance by projecting high-dimensional data to a low-dimensional principal subspace before fitting. However, PCR incurs truncation bias whenever the true regression vector has mass outside the retained principal components (PC). To mitigate the bias, we propose Calibrated Principal Component Regression (CPCR), which first learns a low-variance prior in the PC subspace and then calibrates the model in the original feature space via a centered Tikhonov step. CPCR leverages cross-fitting and controls the truncation bias by softening PCR's hard cutoff. Theoretically, we calculate the out-of-sample risk in the random matrix regime, which shows that CPCR outperforms standard PCR when the regression signal has non-negligible components in low-variance directions. Empirically, CPCR consistently improves prediction across multiple overparameterized problems. The results highlight CPCR's stability and flexibility in modern overparameterized settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。