提出输入空间曲率率λ,量化神经网络决策边界的平滑性。
The Curvature Rate λ: A Scalar Measure of Input-Space Sharpness in Neural Networks
- 定义输入空间曲率率λ,基于高阶导数增长速率。
- 训练中λ可预测演化,用正则化使其更平坦。
- 比SAM更优的置信度校准,且不依赖参数重排。
曲率影响泛化性能、鲁棒性及模型对微小输入扰动的响应稳定性。现有尖锐度度量多在参数空间定义(如海森矩阵特征值),存在计算成本高、对参数重排敏感、功能解释性差等问题。本文提出一种直接定义在输入空间的标量曲率度量——曲率率λ,即高阶输入导数的指数增长速率。实验上,λ通过log ||D^n f||与n的斜率估算。该增长率视角统一了经典解析函数的性质:对解析函数,λ对应收敛半径倒数;对带限信号,λ反映频谱截断。该原理可推广至神经网络,λ能追踪决策边界高频结构的出现。在解析函数及神经网络(Two Moons和MNIST)上的实验表明,λ在训练中可预测演化,并可通过简单基于导数的正则化方法——曲率率正则化(CRR)直接调控。相比尖锐感知最小化(SAM),CRR在保持相近精度的同时,使输入空间几何更平坦,置信度校准更优。通过将曲率建立在微分动态基础上,λ提供了一个紧凑、可解释、参数不变的功能平滑性描述符。
原文摘要 · Abstract (English)
Curvature influences generalization, robustness, and how reliably neural networks respond to small input perturbations. Existing sharpness metrics are typically defined in parameter space (e.g., Hessian eigenvalues) and can be expensive, sensitive to reparameterization, and difficult to interpret in functional terms. We introduce a scalar curvature measure defined directly in input space: the curvature rate λ, given by the exponential growth rate of higher-order input derivatives. Empirically, λ is estimated as the slope of log ||D^n f|| versus n for small n. This growth-rate perspective unifies classical analytic quantities: for analytic functions, λ corresponds to the inverse radius of convergence, and for bandlimited signals, it reflects the spectral cutoff. The same principle extends to neural networks, where λ tracks the emergence of high-frequency structure in the decision boundary. Experiments on analytic functions and neural networks (Two Moons and MNIST) show that λ evolves predictably during training and can be directly shaped using a simple derivative-based regularizer, Curvature Rate Regularization (CRR). Compared to Sharpness-Aware Minimization (SAM), CRR achieves similar accuracy while yielding flatter input-space geometry and improved confidence calibration. By grounding curvature in differentiation dynamics, λ provides a compact, interpretable, and parameterization-invariant descriptor of functional smoothness in learned models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。