arXiv:2602.10613stat.MLcs.LG2026-02被引 3

提出新型主成分降维方法,让高维非参数回归更快更准。

Highly Adaptive Principal Component Regression

  • 用主成分分析降维HAL基函数,大幅降低计算开销。
  • 在高维场景下性能接近最优,且计算速度提升显著。
  • 适合需要快速处理高维数据的统计建模任务。

高度自适应套索(HAL)是一种非参数回归方法,在最小光滑性假设下可实现近乎无维度的收敛速率,但其高维实现因需构造大规模设计矩阵而计算成本高昂。为此,我们提出高度自适应岭(HAR)作为相关岭正则化版本。在此基础上,引入主成分高度自适应套索(PCHAL)与主成分高度自适应岭(PCHAR),通过使用对结果无关的主成分对HAL基进行降维,显著提升计算效率,同时保持与HAL和HAR相当的实证性能。此外,我们提出一种提前停止的梯度下降变体,无需显式设定主成分截断点即可实现平滑谱正则化。最后发现,在特定条件下,HAL核函数等价于布朗运动的协方差函数。

原文摘要 · Abstract (English)

The Highly Adaptive Lasso (HAL) is a nonparametric regression method that achieves almost dimension-free convergence rates under minimal smoothness assumptions, but its implementation can be computationally prohibitive in high dimensions due to the large design matrix it requires. The Highly Adaptive Ridge (HAR) has been proposed as a related ridge-regularized analogue. Building on both procedures, we introduce the Principal Component Highly Adaptive Lasso (PCHAL) and Principal Component Highly Adaptive Ridge (PCHAR). These estimators use an outcome-blind principal-component reduction of the HAL basis, offering substantial computational gains over HAL while achieving empirical performance comparable to HAL and HAR. We also describe an early-stopped gradient descent variant, which provides a convenient form of smooth spectral regularization without explicitly selecting a hard principal-component cutoff. Finally, we uncover that under special circumstances, the HAL kernel is identical to the covariance function of Brownian motion.

非参数回归主成分分析高维统计正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。