arXiv:2503.18219cs.LGmath.FA2025-03被引 6

揭示神经网络与算子学习中的理论到实践的样本效率差距

Theory-to-Practice Gap for Neural Networks and Neural Operators

  • 从泛函空间角度推导出最优学习算法的收敛速率上界
  • 发现有限维下参数复杂度与采样复杂度存在显著差距
  • 首次将该差距扩展至无限维算子学习场景,适用于FNO等模型

本文研究了基于ReLU神经网络和神经算子的学习的采样复杂度。对于属于相关逼近空间的映射,我们推导出任意学习算法在样本数量上的最佳可能收敛速率上界。在有限维情形下,这些上界表明参数复杂度与采样复杂度之间存在差距,即所谓的‘理论到实践差距’。本文在一般的$L^p$设定下实现了对这一差距的统一分析,同时改进了文献中已有的界。此外,基于这些结果,我们将理论到实践的差距推广到了算子学习的无限维情形。我们的结论适用于深度算子网络和基于积分核的神经算子,包括傅里叶神经算子。我们证明,在Bochner $L^p$-范数下,最佳可能的收敛速率被限定在$1/p$阶。

原文摘要 · Abstract (English)

This work studies the sampling complexity of learning with ReLU neural networks and neural operators. For mappings belonging to relevant approximation spaces, we derive upper bounds on the best-possible convergence rate of any learning algorithm, with respect to the number of samples. In the finite-dimensional case, these bounds imply a gap between the parametric and sampling complexities of learning, known as the \emph{theory-to-practice gap}. In this work, a unified treatment of the theory-to-practice gap is achieved in a general $L^p$-setting, while at the same time improving available bounds in the literature. Furthermore, based on these results the theory-to-practice gap is extended to the infinite-dimensional setting of operator learning. Our results apply to Deep Operator Networks and integral kernel-based neural operators, including the Fourier neural operator. We show that the best-possible convergence rate in a Bochner $L^p$-norm is bounded by rates of order $1/p$.

神经算子理论分析收敛速率学习复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。