arXiv:2502.00705cs.LGmath.OC2025-02ICML被引 1

理论证明越宽的神经算子优化越快,解决长期无收敛保障难题。

Optimization for Neural Operators can Benefit from Width

  • 统一框架分析梯度下降优化,揭示宽度影响收敛性
  • 证明两类算子损失满足限制强凸与光滑性,可保证下降
  • 适合关注算子网络理论与架构设计的研究者

神经算子(如深度算子网络DONs和傅里叶神经算子FNOs)直接学习函数空间间的映射,虽具通用逼近能力,但梯度下降(GD)的优化收敛性尚未有理论保障。本文提出统一优化框架,建立对DONs和FNOs的收敛性证明。关键发现:两类模型的损失函数均满足限制强凸性(RSC)与光滑性,从而确保梯度下降过程中损失值递减。值得注意的是,这一性质源于不同模型结构特性。理论推导表明:更宽的网络能促进优化收敛。在典型算子学习任务上进行了实验验证,支持理论结果。

原文摘要 · Abstract (English)

Neural Operators that directly learn mappings between function spaces, such as Deep Operator Networks (DONs) and Fourier Neural Operators (FNOs), have received considerable attention. Despite the universal approximation guarantees for DONs and FNOs, there is currently no optimization convergence guarantee for learning such networks using gradient descent (GD). In this paper, we address this open problem by presenting a unified framework for optimization based on GD and applying it to establish convergence guarantees for both DONs and FNOs. In particular, we show that the losses associated with both of these neural operators satisfy two conditions -- restricted strong convexity (RSC) and smoothness -- that guarantee a decrease on their loss values due to GD. Remarkably, these two conditions are satisfied for each neural operator due to different reasons associated with the architectural differences of the respective models. One takeaway that emerges from the theory is that wider networks should lead to better optimization convergence for both DONs and FNOs. We present empirical results on canonical operator learning problems to support our theoretical results.

神经算子优化理论深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。