arXiv:2606.14954math.FAcs.LG2026-06

揭示深度神经网络权重衰减带来的新型函数空间几何结构

Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks

  • 构建参数正则化下表示成本的统一分析框架
  • 发现深度ReLU网络在权重衰减下形成非凸单位球的拟巴拿赫空间
  • 为深层网络提供理论解释,适合研究模型泛化与结构设计者

我们建立了一个通用框架,用于分析数据拟合方法中由参数空间正则化引发的表示成本。针对任意参数化方法,定义其表示成本和原生函数空间,证明存在性,并确定参数空间与函数空间问题具有相同最小值及极小值可转移的条件。该框架导出表示定理,并恢复经典形式——包括核方法与再生核希尔伯特空间、小波与贝索夫空间、浅层神经网络与变差空间。主要新结果针对带权重衰减的深度为L的前馈ReLU网络:证明表示成本是某拟范数的幂,且在满足一定条件下,当L > 2时,原生空间为具有非凸单位球的拟巴拿赫空间。这些结果揭示了权重衰减诱导的深度相关拟巴拿赫函数空间几何。

原文摘要 · Abstract (English)

We develop a general framework for analyzing representation costs induced by parameter-space regularizers in data-fitting methods. For an arbitrary parametric method, we define its representation cost and native function space, prove existence, and identify conditions under which parameter-space and function-space problems have equal infimal values and minimizers transfer between them. This framework yields representer theorems and recovers classical formulations---including kernel methods and RKHSs, wavelets and Besov spaces, and shallow neural networks and variation spaces---as special cases. Our main new results concern depth-$L$ feedforward ReLU networks with weight-decay regularization. For these networks, we prove that the representation cost is a power of a quasi-seminorm and that, under suitable hypotheses, the native space is a quasi-Banach space with nonconvex unit ball when $L > 2$. These results identify a novel depth-dependent quasi-Banach function-space geometry induced by weight decay.

深度学习函数空间正则化理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。