arXiv:2604.01961cs.LG2026-04被引 1

提出多算子学习的泛化界分析,揭示样本量对模型性能的影响。

Generalization Bounds and Statistical Guarantees for Multi-Task and Multiple Operator Learning with MNO Networks

  • 基于覆盖数方法分析MNO网络的复杂度,建立可解释的泛化边界。
  • 给出在不同采样预算下测试误差的显式上界,明确学习速率与样本量关系。
  • 适用于多任务、多算子场景,尤其适合构建小型通用微分方程模型。

多算子学习关注于学习由算子描述符 α 索引的算子族 {G[α]:U→V}_{α∈W}。训练数据通过分层采样获得:先采样算子实例 α,再每个实例采样输入函数 u,最后在每个输入上采样评估点 x,得到 G[α][u](x) 的噪声观测。尽管近期工作已发展出表达力强的多任务与多算子学习架构及逼近理论的标度律,但定量统计泛化保证仍有限。本文针对可分离模型,特别是多神经算子(MNO)架构,提供基于覆盖数的泛化分析:首先推导出由深度ReLU子网络乘积的线性组合构成的假设类的显式度量熵界,再结合MNO的逼近保证,得到新三元组 (α, u, x) 上期望测试误差的显式逼近-估计权衡。该边界清晰展现了对分层采样预算 (n_α, n_u, n_x) 的依赖,并给出了算子采样预算 n_α 下的显式学习率声明,为跨算子实例的泛化提供了样本复杂度刻画。该结构与架构亦可视为通用求解器或一种‘小’的偏微分方程基础模型,其中三元组体现了一种多模态形式。

原文摘要 · Abstract (English)

Multiple operator learning concerns learning operator families $\{G[α]:U\to V\}_{α\in W}$ indexed by an operator descriptor $α$. Training data are collected hierarchically by sampling operator instances $α$, then input functions $u$ per instance, and finally evaluation points $x$ per input, yielding noisy observations of $G[α][u](x)$. While recent work has developed expressive multi-task and multiple operator learning architectures and approximation-theoretic scaling laws, quantitative statistical generalization guarantees remain limited. We provide a covering-number-based generalization analysis for separable models, focusing on the Multiple Neural Operator (MNO) architecture: we first derive explicit metric-entropy bounds for hypothesis classes given by linear combinations of products of deep ReLU subnetworks, and then combine these complexity bounds with approximation guarantees for MNO to obtain an explicit approximation-estimation tradeoff for the expected test error on new (unseen) triples $(α,u,x)$. The resulting bound makes the dependence on the hierarchical sampling budgets $(n_α,n_u,n_x)$ transparent and yields an explicit learning-rate statement in the operator-sampling budget $n_α$, providing a sample-complexity characterization for generalization across operator instances. The structure and architecture can also be viewed as a general purpose solver or an example of a "small'' PDE foundation model, where the triples are one form of multi-modality.

多算子学习泛化界神经算子统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。