神经网络在组合任务中远超NTK,因前者更适应复杂结构。
A Function-Space Dichotomy for Compositional Learning: Exponential Sub-Optimality of the Neural Tangent Kernel

- 用傅里叶复杂度与架构复杂度区分学习难度,揭示性能差异根源。
- NTK需指数级更多样本(如4^L),而最优网络仅需多项式量级。
- 适合研究深度学习泛化性与核方法局限性的读者。
一个长期观察是:训练后的神经网络在具有组合结构的任务上表现优于其神经正切核(NTK)极限,但何时以及多大程度上存在这种差距尚无定量解释。本文在单位圆上通过两种复杂度测度的二分法给出了答案:目标函数的傅里叶复杂度决定NTK核回归性能,而其架构复杂度则控制深度为L、宽度为w、权重变差范数有界于R的ReLU网络的学习能力。我们首先刻画了架构类𝐶_{L,w,R}的极小极大率,精度达单个因子L:介于Ω(Lw²R²/n)与~O(L²w²R²/n)之间。随后证明,当两种复杂度解耦时,NTK估计器性能会指数级劣于下界:例如对深度为L的迭代锯齿函数,NTK需Ω(4^L)样本,而极小极大下界仅为L的多项式。数值实验验证了这一理论:在带限光滑目标上NTK表现相当或更优,而在超立方体稀疏奇偶模型上,两层网络测试误差比NTK低四到六数量级。该差距本质是函数空间特性,源于核的平滑性偏置与目标组合结构之间的不匹配,而非普适的核与网络现象。
原文摘要 · Abstract (English)
A persistent empirical observation is that trained neural networks outperform their neural tangent kernel (NTK) limit on tasks with compositional structure, yet a quantitative account of $\textbf{when}$ and $\textbf{by how much}$ has been lacking. Working on the unit circle, we give such an account through a dichotomy between two complexity measures of the target: its $\textbf{Fourier complexity}$, which controls NTK kernel regression, and its $\textbf{architectural complexity}$, which controls learning over depth-$L$, width-$w$ ReLU networks with the variation norm of the weights bounded by $R$. We first characterize the minimax rate of the architecture class $\mathcal{C}_{L,w,R}$, pinning it down up to a single factor of $L$: between $Ω(Lw^2R^2/n)$ and $\tilde{O}(L^2w^2R^2/n)$. We then show the NTK estimator sits $\textbf{exponentially}$ above this floor whenever the two complexities decouple: for the depth-$L$ iterated sawtooth, NTK regression needs $Ω(4^L)$ samples while the minimax floor is polynomial in $L$. Numerical experiments confirm the theoretical claims: on bandlimited smooth targets, the NTK is competitive or better, while on the hypercube sparse-parity model, a standard two-layer network beats the NTK by four to six orders of magnitude in test error. The gap is thus a function-space property, a mismatch between the kernel's smoothness bias and the target's compositional structure, rather than a generic kernel-versus-network phenomenon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。