arXiv:2603.15363cs.LGmath.DS2026-03被引 1

深度网络的逼近能力与深度的关系,由流的几何结构决定。

Deep learning and the rate of approximation by flows

  • 用流的几何视角分析深层网络的函数逼近机制。
  • 最小逼近时间对应子芬斯勒流形上的测地距离。
  • 揭示了深度学习与线性近似理论的本质差异。

我们研究了深度残差网络在连续动力系统框架下,其深度对函数逼近能力的影响。该问题可转化为:在给定向量场族 $\ extcal F$ 驱动的流中,逼近一个微分同胚所需的最短时间。我们证明该最小时间可等价为一个子芬斯勒流形上微分同胚的测地距离,其局部几何由涉及 $\ extcal F$ 的变分原理刻画。这一结果将目标函数的学习效率与其与网络架构的兼容性联系起来。此外,研究揭示深度学习中通过组合或动态逼近函数的机制,从根本上区别于线性逼近理论——后者依赖线性空间和范数估计,而前者基于流形与测地距离。

原文摘要 · Abstract (English)

We investigate the dependence of the approximation capacity of deep residual networks on its depth in a continuous dynamical systems setting. This can be formulated as the general problem of quantifying the minimal time-horizon required to approximate a diffeomorphism by flows driven by a given family $\mathcal F$ of vector fields. We show that this minimal time can be identified as a geodesic distance on a sub-Finsler manifold of diffeomorphisms, where the local geometry is characterised by a variational principle involving $\mathcal F$. This connects the learning efficiency of target relationships to their compatibility with the learning architectural choice. Further, the results suggest that the key approximation mechanism in deep learning, namely the approximation of functions by composition or dynamics, differs in a fundamental way from linear approximation theory, where linear spaces and norm-based rate estimates are replaced by manifolds and geodesic distances.

深度学习流模型逼近理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。