arXiv:2602.21910cs.LGcs.NA2026-02

剖析深度算子网络误差来源,发现分支与模式误差主导性能瓶颈。

The Error of Deep Operator Networks Is the Sum of Its Parts: Branch-Trunk and Mode Error Decompositions

  • 分解误差为分支网络与模式学习部分,揭示主因
  • 低频模式更易学习,小奇异值对应模式泛化差
  • 共享分支结构提升小模式泛化,宽度增大缓解模式耦合

算子学习有望通过学习微分方程的解算子,显著加速多查询任务如设计优化与不确定性量化。尽管深层算子网络(DeepONets)具备通用逼近性质,实际中常表现有限精度与泛化能力,制约其应用。本文分析经典DeepONet架构的性能局限:当内部维度足够大时,近似误差主要由分支网络决定;且所学的托辊基函数可被经典基函数替代而性能无显著下降。为此构建改进型DeepONet,用训练解矩阵的左奇异向量替换托辊网络。结果显示:对KdV与Burgers方程,分支网络存在谱偏差,低频主导模式学习更优;奇异值加权与优化器忽略小模式共同导致分支误差集中于中高频模式;标准共享分支结构比独立计算系数的堆叠架构更利于小模式泛化;参数空间中的模式间有害耦合随网络宽度增加而减弱。

原文摘要 · Abstract (English)

Operator learning has the potential to strongly impact scientific computing by learning solution operators for differential equations, potentially accelerating multi-query tasks such as design optimization and uncertainty quantification by orders of magnitude. Despite proven universal approximation properties, deep operator networks (DeepONets) often exhibit limited accuracy and generalization in practice, which hinders their adoption. Understanding these limitations is therefore crucial for further advancing the approach. This work analyzes performance limitations of the classical DeepONet architecture. It is shown that the approximation error is dominated by the branch network when the internal dimension is sufficiently large, and that the learned trunk basis can often be replaced by classical basis functions without a significant impact on performance. To investigate this further, a modified DeepONet is constructed in which the trunk network is replaced by the left singular vectors of the training solution matrix. This modification yields several key insights. First, for examples involving the KdV and Burgers equations, a spectral bias in the branch network is observed, with coefficients of dominant, low-frequency modes learned more effectively. Second, through the interplay of the singular-value weighting and the optimizer's neglect of small modes, the branch error is dominated by modes with large and intermediate singular values. Third, using a shared branch network for all mode coefficients, as in the standard architecture, improves generalization of small modes compared to a stacked architecture in which coefficients are computed separately. Finally, detrimental coupling between modes in parameter space is identified, which weakens with increasing network width.

算子学习深度网络误差分析谱偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。