深度网络首次实现流形上函数与导数的同步逼近,突破维度诅咒。
Expressive Power of Deep Networks on Manifolds: Simultaneous Approximation
- 用固定深度的ReLU^{k-1}网络实现流形上函数及导数的联合逼近
- 参数量仅随ε^{-d/(k-s)}增长,依赖内在维数d而非外在维度
- 理论证明接近最优,适用于复杂几何上的微分方程求解
科学机器学习中求解复杂域上的偏微分方程(PDE)面临关键挑战:曲面几何使函数及其导数的逼近变得困难。本文建立了首个关于流形上深度神经网络的同步逼近理论。我们证明,具有有界权重的常数深度ReLU^{k-1}网络(对控制泛化误差至关重要)可将Sobolev空间$\/mathcal{W}_p^{k}(\/mathcal{M}^d)$中的任意函数,在$\/mathcal{W}_p^{s}(\/mathcal{M}^d)$范数下逼近至误差ε,其中k≥3,s<k,所需非零参数量为$\/mathcal{O}(\\varepsilon^{-d/(k-s)})$,该速率克服了维度诅咒,仅依赖于内在维数d。结果可推广至Hölder-Zygmund空间。我们进一步给出了匹配的下界,证明构造近乎最优,参数量相差仅对数因子。下界证明引入了网络高阶导数类的Vapnik-Chervonenkis维数和伪维数的新估计。这些复杂度界限为涉及导数的流形上PDE学习提供了理论基础。分析揭示网络架构利用稀疏结构高效捕捉流形的低维几何特性。最后,数值实验验证了理论结果。
原文摘要 · Abstract (English)
A key challenge in scientific machine learning is solving partial differential equations (PDEs) on complex domains, where the curved geometry complicates the approximation of functions and their derivatives required by differential operators. This paper establishes the first simultaneous approximation theory for deep neural networks on manifolds. We prove that a constant-depth $\mathrm{ReLU}^{k-1}$ network with bounded weights--a property that plays a crucial role in controlling generalization error--can approximate any function in the Sobolev space $\mathcal{W}_p^{k}(\mathcal{M}^d)$ to an error of $\varepsilon$ in the $\mathcal{W}_p^{s}(\mathcal{M}^d)$ norm, for $k\geq 3$ and $s<k$, using $\mathcal{O}(\varepsilon^{-d/(k-s)})$ nonzero parameters, a rate that overcomes the curse of dimensionality by depending only on the intrinsic dimension $d$. These results readily extend to functions in Hölder-Zygmund spaces. We complement this result with a matching lower bound, proving our construction is nearly optimal by showing the required number of parameters matches up to a logarithmic factor. Our proof of the lower bound introduces novel estimates for the Vapnik-Chervonenkis dimension and pseudo-dimension of the network's high-order derivative classes. These complexity bounds provide a theoretical cornerstone for learning PDEs on manifolds involving derivatives. Our analysis reveals that the network architecture leverages a sparse structure to efficiently exploit the manifold's low-dimensional geometry. Finally, we corroborate our theoretical findings with numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。