arXiv:2603.17785math.OCcs.AI2026-03

揭示无限宽神经网络稀疏性的几何根源,为理解模型简化提供理论支撑。

A Dual Certificate Approach to Sparsity in Infinite-Width Shallow Neural Networks

  • 利用对偶证书的分段线性结构,从数据诱导的超平面划分出发分析解的稀疏性。
  • 证明解的支撑集有限,其数量由数据几何决定,且在低噪声小正则下保持不变。
  • 适用于关注理论解释、稀疏建模与泛化性能的机器学习研究者。

本文研究无限宽浅层ReLU神经网络在总变差(TV)正则化下的训练问题,将其形式化为单位球面上测度上的凸优化问题。通过利用TV正则化优化问题的对偶理论,建立了训练解稀疏性的严格保证。分析进一步刻画了在低噪声和小正则化参数条件下稀疏性如何保持。关键观察是:对于ReLU激活函数,对应的对偶证书在权重空间中是分段线性的,其线性区域(称为对偶区域)由数据诱导的超平面排列决定。利用这一结构,证明了在每个对偶区域内,对偶证书最多只有一个极值点。因此,任意最小化器的支撑集有限,且其基数可被仅依赖于数据诱导超平面排列几何的常数上界控制。随后,我们研究了保证稀疏解唯一性的充分条件。最后,在对偶区域边界上对偶证书满足合适非退化条件下,我们证明:当标签噪声较低且正则化参数较小时,训练解仍保持稀疏,具有相同的狄拉克峰数量,其位置与幅值收敛;若位置位于对偶区域内部,则收敛速率与噪声和正则化参数呈线性关系。

原文摘要 · Abstract (English)

In this paper, we study total variation (TV)-regularized training of infinite-width shallow ReLU neural networks, formulated as a convex optimization problem over measures on the unit sphere. Our approach leverages the duality theory of TV-regularized optimization problems to establish rigorous guarantees on the sparsity of the solutions to the training problem. Our analysis further characterizes how and when this sparsity persists in a low noise regime and for small regularization parameter. The key observation that motivates our analysis is that, for ReLU activations, the associated dual certificate is piecewise linear in the weight space. Its linearity regions, which we name dual regions, are determined by the activation patterns of the data via the induced hyperplane arrangement. Taking advantage of this structure, we prove that, on each dual region, the dual certificate admits at most one extreme value. As a consequence, the support of any minimizer is finite, and its cardinality can be bounded from above by a constant depending only on the geometry of the data-induced hyperplane arrangement. Then, we further investigate sufficient conditions ensuring uniqueness of such sparse solution. Finally, under a suitable non-degeneracy condition on the dual certificate along the boundaries of the dual regions, we prove that in the presence of low label noise and for small regularization parameter, solutions to the training problem remain sparse with the same number of Dirac deltas. Additionally, their location and the amplitudes converge, and, in case the locations lie in the interior of a dual region, the convergence happens with a rate that depends linearly on the noise and the regularization parameter.

神经网络理论稀疏性凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。