arXiv:2602.20921cs.LG2026-02

从动力系统视角揭示深度残差网络泛化能力的内在机制

On the Generalization Behavior of Deep Residual Networks From a Dynamical System Perspective

  • 用动力系统方法结合径向复杂度分析泛化误差
  • 证明了样本数 $S$ 增加时误差以 $O(1/\sqrt{S})$ 下降
  • 适用于离散与连续两种残差网络,统一理解其泛化行为

深度神经网络(DNN)在机器学习中取得显著进展,模型深度在其成功中起关键作用。动力系统建模方法近年来成为有力框架,为DNN的结构和学习行为提供新数学洞见。本文通过结合径向复杂度、动力系统流映射及深层极限下残差网络(ResNet)的收敛性,建立了离散与连续时间残差网络的泛化误差界。所得边界关于训练样本数 $S$ 为 $O(1/\sqrt{S})$,并包含依赖结构的负项,在较弱假设下实现深度无关与渐近泛化界。这些结果统一了离散与连续时间残差网络的泛化理解,弥合了两者在样本复杂度阶数与假设条件上的差距。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) have significantly advanced machine learning, with model depth playing a central role in their successes. The dynamical system modeling approach has recently emerged as a powerful framework, offering new mathematical insights into the structure and learning behavior of DNNs. In this work, we establish generalization error bounds for both discrete- and continuous-time residual networks (ResNets) by combining Rademacher complexity, flow maps of dynamical systems, and the convergence behavior of ResNets in the deep-layer limit. The resulting bounds are of order $O(1/\sqrt{S})$ with respect to the number of training samples $S$, and include a structure-dependent negative term, yielding depth-uniform and asymptotic generalization bounds under milder assumptions. These findings provide a unified understanding of generalization across both discrete- and continuous-time ResNets, helping to close the gap in both the order of sample complexity and assumptions between the discrete- and continuous-time settings.

深度学习残差网络泛化分析动力系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。