arXiv:2501.03017cs.LG2025-01被引 9

揭示ReLU神经网络凸性的充要条件,打破仅用ICNN保证凸性的局限。

Convexity in ReLU Neural Networks: beyond ICNNs?

  • 提出ReLU网络凸性的权重与激活乘积判据,适用于路径提升框架架构。
  • 证明单隐层网络的凸函数可由同结构ICNN实现,多层则不可。
  • 给出可精确验证大规模分段线性网络凸性的数值方法,适合理论研究者。

凸函数及其梯度在数学成像中至关重要,涵盖近端优化到最优传输。深度学习的成功促使人们用可学习的神经网络替代固定函数或算子。尽管这些方法经验上表现优越,但建立严格保证通常需对网络结构施加约束,尤其是凸性。最常用的方法是输入凸神经网络(ICNN)。为探索ICNN的表达能力,本文给出了ReLU网络具有凸性的必要充分条件,其基于权重与激活的乘积,且适用于路径提升框架中的任意架构。作为具体应用,我们深入研究了一层和两层隐藏层网络:证明任一单隐层ReLU网络实现的凸函数均可由同结构的ICNN表示;但该性质在更多层数下不再成立。最后,我们提供一种数值算法,可精确检查具有大量仿射区域的ReLU网络的凸性。

原文摘要 · Abstract (English)

Convex functions and their gradients play a critical role in mathematical imaging, from proximal optimization to Optimal Transport. The successes of deep learning has led many to use learning-based methods, where fixed functions or operators are replaced by learned neural networks. Regardless of their empirical superiority, establishing rigorous guarantees for these methods often requires to impose structural constraints on neural architectures, in particular convexity. The most popular way to do so is to use so-called Input Convex Neural Networks (ICNNs). In order to explore the expressivity of ICNNs, we provide necessary and sufficient conditions for a ReLU neural network to be convex. Such characterizations are based on product of weights and activations, and write nicely for any architecture in the path-lifting framework. As particular applications, we study our characterizations in depth for 1 and 2-hidden-layer neural networks: we show that every convex function implemented by a 1-hidden-layer ReLU network can be also expressed by an ICNN with the same architecture; however this property no longer holds with more layers. Finally, we provide a numerical procedure that allows an exact check of convexity for ReLU neural networks with a large number of affine regions.

凸性分析ReLU网络ICNN理论保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。