揭示浅层ReLU网络损失曲面连通性与逼近率的深层关联,给出可证明的优化路径。
From Approximation Rates to Loss-Landscape Barrier Decay in Shallow ReLU Networks
- 用光滑压缩泛函与一阶扰动构建损失一致路径,替代传统二次估计。
- 在n≥2维时,损失曲面厚化率达O(m^{-1/(n-1)}),一维情形对任意m≥4精确连通。
- 理论揭示正则化逼近误差衰减速率与障碍物衰减速率的对应关系,适用于小样本优化分析。
研究带约束第一层权重和输出层ℓ₁惩罚的一层ReLU网络子水平集的路径连通性。假设数据项在标量逻辑值上为凸且全局Lipschitz连续。首先构造有限宽度路径,连接同一子水平集内任意两点,路径受损失一致压缩泛函与一阶扰动项控制。证明中以直接Lipschitz界替代Freeman–Bruna机制中的二次扰动估计。利用正齐次性,在与惩罚相容的方向上,每个活跃原子从单位球内单调移至单位球面,同时输出系数减小。通过球面覆盖与聚类合并,得到n≥2时固定水平的厚化率为O(m^{-1/(n-1)});一维双射字典情形下,对任意m≥4实现精确连通。进一步证明正则化逼近值满足e(l)−e_∞=O(l^{-1/2})。更一般地,若逼近速率为O(l^{-s}),则对应近最优障碍衰减率为O(m^{-s/((n−1)s+1)});在现有假设下,具体得O(m^{-1/(n+1)})。定理对齐的有限分布实验补充分析:主Huber实验在宽度m≥16时,720组配对最大认证上界间隙达1.66×10⁻⁵;匹配的二元交叉熵重跑及720端点密集表示压力测试验证损失鲁棒性与活跃聚类合并机制。
原文摘要 · Abstract (English)
We study pathwise connectivity of sublevel sets for one-hidden-layer ReLU networks with constrained first-layer weights and an $\ell_1$ penalty on the output layer. The data term is assumed convex and globally Lipschitz in the scalar logit. We first give a finite-width construction that connects any two points of a common sublevel through a path controlled by a loss-consistent compression functional and a first-order perturbation term. The proof replaces the quadratic perturbation estimate in the Freeman--Bruna mechanism by a direct Lipschitz bound. Positive homogeneity is then used in a direction that is compatible with the penalty: every active atom is moved monotonically from the unit ball to the unit sphere while its output coefficient is reduced. Sphere covering and cluster merging consequently give $O(m^{-1/(n-1)})$ fixed-level thickening for $n\ge2$, while the one-dimensional two-ray dictionary gives exact connectivity for every $m\ge4$. We also prove internally that the regularized approximation values satisfy $e(l)-e_\infty=O(l^{-1/2})$. More generally, a rate $O(l^{-s})$ transfers to a near-optimal barrier rate $O(m^{-s/((n-1)s+1)})$; under the standing assumptions, this yields the explicit rate $O(m^{-1/(n+1)})$. A theorem-aligned finite-distribution experiment complements the analysis. The primary Huber run yields a maximal best certified upper gap $1.66\times10^{-5}$ over 720 recorded pairs at widths $m\ge16$; a matched binary-cross-entropy rerun and a 720-endpoint dense-representation stress test probe loss robustness and the active cluster-merging mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。