揭示过参数线性模型中参数范数随数据量变化的精确规律
Closed-form $\ell_r$ norm scaling with data for overparameterized linear regression and diagonal linear networks under $\ell_p$ bias
- 通过双射分析揭示信号尖峰与噪声主体的竞争机制
- 发现数据量转折点n_⋆和范数阈值r_⋆=2(p−1)的闭式解
- 适用于理解模型泛化能力中的范数敏感性,适合理论研究者
针对具有各向同性高斯设计的过参数化线性回归及最小ℓ_p插值(p∈(1,2]),本文给出参数范数族{‖ŵ_p‖_r}_{r∈[1,p]}在样本量下的统一、高概率刻画。通过简洁的对偶射线分析,揭示了XᵀY中信号尖峰与零坐标主体之间的竞争,得到:(i) 数据依赖的转折点n_⋆(“拐点”)的闭式预测;(ii) 一个通用阈值r_⋆=2(p−1),将‖ŵ_p‖_r分为平台区与持续增长区,且后者有明确指数。该统一解完整解析了ℓ_r范数在r∈[1,p]范围内的标度行为,清晰展示哪些范数随样本量饱和,哪些继续增长。进一步研究对角线性网络(DLNs)在梯度下降下的训练,通过初始化尺度α校准有效指数p_eff(α),实验证明其继承相同拐点与阈值规律,建立起显式与隐式偏差间的预测桥梁。由于众多泛化代理指标依赖‖ŵ_p‖_r,结果表明其预测力高度依赖所选ℓ_r范数。
原文摘要 · Abstract (English)
For overparameterized linear regression with isotropic Gaussian design and minimum-$\ell_p$ interpolator $p\in(1,2]$, we give a unified, high-probability characterization for the scaling of the family of parameter norms $ \\{ \lVert \widehat{w_p} \rVert_r \\}_{r \in [1,p]} $ with sample size. We solve this basic, but unresolved question through a simple dual-ray analysis, which reveals a competition between a signal *spike* and a *bulk* of null coordinates in $X^\top Y$, yielding closed-form predictions for (i) a data-dependent transition $n_\star$ (the "elbow"), and (ii) a universal threshold $r_\star=2(p-1)$ that separates $\lVert \widehat{w_p} \rVert_r$'s which plateau from those that continue to grow with an explicit exponent. This unified solution resolves the scaling of *all* $\ell_r$ norms within the family $r\in [1,p]$ under $\ell_p$-biased interpolation, and explains in one picture which norms saturate and which increase as $n$ grows. We then study diagonal linear networks (DLNs) trained by gradient descent. By calibrating the initialization scale $α$ to an effective $p_{\mathrm{eff}}(α)$ via the DLN separable potential, we show empirically that DLNs inherit the same elbow/threshold laws, providing a predictive bridge between explicit and implicit bias. Given that many generalization proxies depend on $\lVert \widehat {w_p} \rVert_r$, our results suggest that their predictive power will depend sensitively on which $l_r$ norm is used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。