揭示深度算子网络的理论缩放规律,解释模型与数据规模如何影响性能。
Neural Scaling Laws of Deep ReLU and Deep Operator Network: A Theoretical Study
- 通过分析近似与泛化误差,建立算子网络缩放定律的理论框架。
- 发现模型大小和训练数据量共同决定误差上限,且低维结构可提升精度。
- 结果适用于深度ReLU网络,为机器学习模型设计提供理论依据。
神经缩放定律在深度神经网络性能中起关键作用,并已在多种任务中被观察到。然而,对其完整的理论框架仍不充分。本文研究深度算子网络的缩放定律,聚焦于Chen和Chen风格的架构,这类方法包括流行的Deep Operator Network(DeepONet),通过可学习基函数与依赖输入函数的系数的线性组合来近似输出函数。我们建立理论框架,量化其近似与泛化误差,阐明误差与网络规模、训练数据量之间的关系。此外,针对输入函数具有低维结构的情形,推导出更紧的误差界。这些结果也适用于深度ReLU网络及其他类似结构。本工作部分解释了算子学习中的缩放现象,为相关应用提供了理论基础。
原文摘要 · Abstract (English)
Neural scaling laws play a pivotal role in the performance of deep neural networks and have been observed in a wide range of tasks. However, a complete theoretical framework for understanding these scaling laws remains underdeveloped. In this paper, we explore the neural scaling laws for deep operator networks, which involve learning mappings between function spaces, with a focus on the Chen and Chen style architecture. These approaches, which include the popular Deep Operator Network (DeepONet), approximate the output functions using a linear combination of learnable basis functions and coefficients that depend on the input functions. We establish a theoretical framework to quantify the neural scaling laws by analyzing its approximation and generalization errors. We articulate the relationship between the approximation and generalization errors of deep operator networks and key factors such as network model size and training data size. Moreover, we address cases where input functions exhibit low-dimensional structures, allowing us to derive tighter error bounds. These results also hold for deep ReLU networks and other similar structures. Our results offer a partial explanation of the neural scaling laws in operator learning and provide a theoretical foundation for their applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。