arXiv:2606.08799stat.MLcs.LG2026-06

通过学习到的特征几何,揭示非线性最小二乘模型泛化能力的本质。

Generalization in Nonlinear Least Squares via Learned Feature Geometry

论文配图:Generalization in Nonlinear Least Squares via Learned Feature Geometry
图 1 · 摘自论文原文
  • 基于梯度模型几何与残差曲率,定义数据依赖的有效维度。
  • 在流形数据下,泛化误差界与内在维数相关而非参数量。
  • 适用于神经网络训练后的稳定性分析,推导简洁且可解释。

我们通过平均算法稳定性研究了岭正则化非线性最小二乘模型的泛化能力,推导出局部极小值的误差界,其依赖于数据相关的有效维度,该维度由训练参数处的经验雅可比格拉姆矩阵和残差曲率项反映。在线性情形下,曲率项消失,结果恢复经典雅可比核协方差的有效维度,但评估点为训练后参数而非初始化点,区别于神经切线核分析。我们进一步通过梯度特征的覆盖复杂度界定该有效维度,得到依赖于学习几何而非参数数量的泛化保证。对于流形支持的数据和分段Lipschitz雅可比矩阵,边界随内在维数增长;对单隐层ReLU网络,可通过激活稳定区域的数量显式刻画机制。在合成流形、聚类分布及基准数据集上的实验展示了训练雅可比压缩现象、残差曲率线性化的紧致性,以及稳定性界与实际泛化差距的一致性。关键优势在于推导简单,基于强对数凹噪声下的Brascamp-Lieb不等式从基本原理出发。

原文摘要 · Abstract (English)

We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-dependent effective dimension that reflects the geometry of the gradient model at the trained parameters, through the empirical Jacobian Gram matrix and a residual-curvature term. In the linear case, where the curvature term vanishes, this recovers the classical effective dimension of the Jacobian kernel covariance, but evaluated at the trained model rather than at initialization as is typical in neural tangent kernel analyses. We further bound this effective dimension via covering complexity of the gradient features, leading to guarantees that depend on learned geometry rather than parameter count. In particular, for manifold-supported data and piecewise Lipschitz Jacobians, the bounds scale with intrinsic dimension, while for one-hidden-layer ReLU networks, the mechanism can be made explicit through counts of activation-stable regions. Experiments on synthetic manifolds, clustered distributions, and benchmark datasets illustrate trained-Jacobian compression, the tightness of the residual-curvature linearization, and agreement between the stability bound and observed generalization gaps. A key feature of our bounds is the simplicity of their derivation, which follows from first principles using the Brascamp-Lieb inequality under strongly log-concave noise.

泛化理论非线性拟合特征几何稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。