ReLU网络的稳定性分析不能依赖光滑性假设,否则会失效。
Why Smooth Stability Assumptions Fail for ReLU Learning
- 证明全局光滑性代理(如梯度Lipschitz)在ReLU网络中无法成立
- 给出反例:即使训练轨迹稳定,经典稳定性界仍不成立
- 提出最小广义导数条件,为非光滑稳定性框架奠基
现代学习系统的稳定性分析常基于光滑性假设,但这类假设在ReLU型非线性中被破坏。本文通过构造最小障碍,证明无论在何种简单设置下,对ReLU网络而言,不存在全局有效的光滑性稳定性代理(如梯度Lipschitz或海森控制)。即使训练轨迹表现出经验上的稳定性,经典稳定性界依然失效。我们给出具体反例,并识别出可使稳定性结论恢复的最小广义导数条件。结果澄清了为何光滑近似可能产生误导,推动了非光滑感知的稳定性框架发展。
原文摘要 · Abstract (English)
Stability analyses of modern learning systems are frequently derived under smoothness assumptions that are violated by ReLU-type nonlinearities. In this note, we isolate a minimal obstruction by showing that no uniform smoothness-based stability proxy such as gradient Lipschitzness or Hessian control can hold globally for ReLU networks, even in simple settings where training trajectories appear empirically stable. We give a concrete counterexample demonstrating the failure of classical stability bounds and identify a minimal generalized derivative condition under which stability statements can be meaningfully restored. The result clarifies why smooth approximations of ReLU can be misleading and motivates nonsmooth-aware stability frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。