arXiv:2608.15982cs.LG2026-08

用算子理论为多任务深度学习提供可分离的泛化界,揭示输出耦合与层间结构的关系。

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

  • 将网络层建模为向量值再生核希尔伯特空间上的Koopman算子,通过算子范数分析泛化能力。
  • 在不同假设空间下得到紧致的泛化界,其中输出耦合由任务矩阵迹控制,层间因子含权重和激活导数的平方根。
  • 适用于多任务学习中的共享算子学习与目标迁移场景,对模型设计有理论指导意义。

我们通过将网络层表示为向量值再生核希尔伯特空间上的Koopman复合算子,建立了深度多输出函数类的算子理论泛化边界。在向量值Sobolev再生核希尔伯特空间中,推导了可逆且宽度扩展的单射架构的Rademacher复杂度界。估计结果将输出耦合贡献(由任务矩阵的迹表示)与逐层算子范数、Sobolev符号比、行列式因子及限制常数分离开来。随后分析了一维布朗运动/卡姆登-马丁情形,利用向量值布朗再生核希尔伯特空间的精确锚定导数范数刻画,获得了保持定义域的标量线性映射和锚定微分同胚激活的逐层边界;对应的因子分别以$|W_l|^{1/2}$和$\\

原文摘要 · Abstract (English)

We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures. The estimates separate the output-coupling contribution, represented by the trace of the task matrix, from the layerwise operator norms, Sobolev symbol ratios, determinant factors, and restriction constants generated by the linear maps. We then analyze a distinct one-dimensional Brownian/Cameron--Martin regime. Using the exact anchored derivative-norm characterization of the vector-valued Brownian RKHS, we obtain layerwise bounds for domain-preserving scalar linear maps and anchored diffeomorphic activations; the corresponding factors scale as $|W_l|^{1/2}$ and $\|σ_l'\|_\infty^{1/2}$, respectively, and do not involve Sobolev smoothness exponents. Because the Sobolev and Brownian results concern different hypothesis spaces, neither is asserted to dominate the other uniformly. We additionally formulate shared operator learning across tasks, prove a finite-rank representer theorem, derive the exact finite-dimensional problem for squared loss, and establish a target-transfer bound when the learned operator is obtained independently of the target sample. Synthetic and MNIST studies examine stabilized Sobolev-inspired and Brownian-inspired complexity proxies; these empirical proxies are not evaluations of the proved bounds for rank-deficient architectures.

多任务学习泛化界算子理论深度学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。