arXiv:2605.12648cs.LGstat.ML2026-05被引 2

首次分析带相关噪声的私有训练在KAN上的泛化性能,贴近实际训练场景。

Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise

  • 引入辅助非投影动态,处理相关噪声与投影带来的数学障碍
  • 证明了在非凸条件下仍能获得可量化的泛化误差界
  • 适合关注差分隐私与神经网络理论的科研人员

本文首次建立了基于小批量SGD与梯度裁剪的柯尔莫哥洛夫-阿诺德网络(KAN)的总体风险边界,涵盖非私有SGD及具有高斯扰动的差分隐私SGD(DP-SGD),其噪声机制在独立与时间相关噪声间插值。该设定比以往的KAN理论更贴近实际:训练采用现代网络的标准小批量SGD,而非全批量梯度下降(GD);且实证表明相关噪声机制在隐私-效用权衡上优于独立噪声机制。结果覆盖了Wang等人(2026)针对全批量GD和独立噪声DP-GD的结论,同时在固定第二层情形下给出更紧的特例边界。核心技术在于提出一种新分析路径,适用于非凸场景下的相关噪声差分隐私训练。时间依赖破坏了标准单步SGD中条件中心结构,投影步骤也阻碍了相关扰动的精确抵消。通过引入一个辅助未投影动力系统、吸收当前噪声扰动的偏移迭代,以及高概率的投影不活跃性验证,结合稳定性泛化论证,最终得到总体风险边界。据我们所知,这是首个对非凸学习中除凸情况外的相关噪声机制进行优化与泛化分析的研究,尤其针对神经网络。

原文摘要 · Abstract (English)

We establish the first population risk bounds for Kolmogorov-Arnold Networks (KANs) trained by mini-batch SGD with gradient clipping, covering non-private SGD as well as differentially private SGD (DP-SGD) with Gaussian perturbations that interpolate between independent and temporally correlated noise. This setting is substantially closer to practice than prior KAN theory along two axes: training is by mini-batch SGD, the standard recipe for modern networks, rather than full-batch gradient descent (GD); and correlated-noise mechanisms have empirically shown a more favorable privacy-utility tradeoff than independent-noise mechanisms. Our results cover the corresponding full-batch GD and independent-noise DP-GD results for KANs by Wang et al. (2026), while yielding sharper fixed-second-layer specializations. The technical core is a new analysis route for correlated-noise DP training in the non-convex regime. Temporal dependence breaks the conditional-centering structure underlying standard one-step SGD arguments, and the projection step obstructs the exact cancellation structure of correlated perturbations. We address these difficulties through an auxiliary unprojected dynamics, a shifted iterate that absorbs the current noise perturbation, and a high-probability bootstrap certifying projection inactivity. Combining this optimization analysis with a stability-based generalization argument yields the stated population risk bounds. To the best of our knowledge, this is the first optimization and population risk analysis of a correlated-noise mechanism for DP training beyond convex learning, in particular for neural networks.

差分隐私神经网络泛化界相关噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。