arXiv:2608.28564stat.MLcs.LG2026-08

揭示输入几何如何决定核模型的泛化性能

Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy

  • 分析幂律各向异性下核岭回归的谱与误差渐近行为
  • 弱各向异性时方差峰值随样本量变化,强时趋于恒定
  • 目标对齐主方向时偏差在非整数样本复杂度下降

我们研究了在各向异性高斯数据下的核岭回归,其中输入协方差以幂律衰减,指数为 α≥0,适用于多项式内积核。在高维极限情形 n=Θ(d^κ) 下,推导出核谱和泛化误差的渐近精确表达式,揭示了各向异性如何重塑学习曲线。当弱各向异性(0<α<1)时,问题仍保持有效高维特性,部分保留各向同性情形特征:方差在整数样本复杂度 κ∈ℕ 处出现峰值,但随 α 增大而减弱;而对于与数据主方向高度对齐的目标,偏差在分数样本复杂度处下降,使偏差转变与插值峰值解耦。当强各向异性(α>1)时,有效维度恒定,方差不再依赖样本量,无论是无正则化(ridgeless)还是固定正则化下均呈现平台或显式衰减速率。偏差发生尖锐跃变,由目标衰减速率决定:低于阈值时学习突变而非渐进;高于阈值时偏差按幂律衰减,恢复经典源和容量率。最后将结果特化到单指标目标,表明指标与数据主方向的对齐程度决定了各向异性对学习的影响。总体而言,这些结果阐明了输入几何如何塑造核特征并根本影响其泛化性质。

原文摘要 · Abstract (English)

We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $α\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime $n=Θ(d^κ)$, revealing how anisotropy reshapes the learning curves. For weak anisotropy ($0<α<1$), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities $κ\in\mathbb{N}$, but these peaks are progressively damped as $α$ grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transitions from the interpolation peaks. For strong anisotropy ($α> 1$), the effective dimension of the problem is constant, and the variance stops depending on sample size altogether, plateauing under ridgeless interpolation or vanishing at an explicit rate under fixed ridge penalty. The bias undergoes a sharp transition governed by the target's decay rate: below a threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law that recovers the classical source and capacity rates. We finally specialize these results to single-index targets, showing how the alignment of the index with the data's principal directions determines the effect of anisotropy on learning. Together, our results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.

核方法泛化误差高维统计各向异性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。