arXiv:2608.25513stat.MLcs.LG2026-08

提出一种自适应正则化方法,无需先验知识即可实现最优学习率。

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

论文配图:Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees
图 1 · 摘自论文原文
  • 基于邻近早停规则,仅比较相邻正则化参数的估计器。
  • 在标准条件下达到近似最优的多项式学习率,含对数因子。
  • 适用于模型正确设定与部分误设场景,计算高效可扩展。

随机特征方法为核岭回归(KRR)提供了一种可扩展的近似,但实现最优学习率所需的正则化参数依赖于未知的光滑性和容量参数。本文提出一种针对随机特征核岭回归(KRR-RF)的邻近早停规则,该方法在逆正则化参数的均匀网格上进行,仅比较相邻估计器,减少了标准全对比较型Lepskii方法的差异比较次数。邻近差异及其经验复杂度项可在随机特征空间中直接计算,无需构造精确的核格矩阵。我们建立了邻近KRR-RF估计器的高概率比较界,并证明在标准源条件与容量条件,以及合适的网格和随机特征预算条件下,所选估计器能达到接近最优的多项式学习率(含对数因子)。该结果允许在不需预先知道源指数与容量指数的情况下选择正则化参数,涵盖完全正确设定与部分误设情形。分析基于一个经验随机特征有效维度,将可观测的停止阈值与随机特征模型的总体复杂度关联起来。模拟与真实数据实验展示了该方法在预测性能与计算行为上优于标准调参过程。

原文摘要 · Abstract (English)

Random feature methods provide a scalable approximation to kernel ridge regression (KRR), but the regularization parameter that yields the oracle learning rate depends on unknown smoothness and capacity parameters. In this work, we propose a neighboring early-stopping rule for adaptive regularization in KRR with random features (KRR-RF). The method uses a grid that is uniform in inverse regularization and compares only adjacent estimators, reducing the number of discrepancy comparisons relative to standard all-pairs Lepskii-type procedures. Both the neighboring discrepancy and its empirical complexity term can be computed directly in the random feature space, without constructing the exact kernel Gram matrix. We establish a high-probability comparison bound for neighboring KRR-RF estimators and show that, under standard source and capacity conditions together with suitable grid and random feature budget conditions, the selected estimator attains the oracle polynomial learning rate up to logarithmic factors. The result allows the regularization parameter to be selected without prior knowledge of the source and capacity exponents and covers both well-specified and partially misspecified regimes. Our analysis is based on an empirical random feature effective dimension that connects the observable stopping threshold with the population complexity of the random feature model. Simulation and real-data experiments illustrate the prediction performance and computational behavior of the proposed method in comparison with standard tuning procedures.

机器学习正则化随机特征自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。