提出统一方法,高效凸化多种神经网络激活函数组合。
Tightening convex relaxations of trained neural networks: a unified approach for convex and S-shaped activations
- 基于递归公式,统一处理凸型与S形激活函数的凸化。
- 可高效计算分离超平面或判断其不存在性。
- 适用于非多面体情形,对优化建模有实用价值。
训练后神经网络的非凸性给其融入优化模型带来了显著挑战。Anderson等(2020)提出了一个框架,用于获取分段线性凸激活函数与仿射函数复合后图像的凸包,从而有效凸化如ReLU及其前接仿射变换。本文在此基础上,提出一种递归公式,适用于广泛激活函数类别——凸型或'S形',可高效计算复合函数的紧凸化,并在多种场景下(包括非多面体情形)用于确定分离超平面是否存在或计算其存在性。
原文摘要 · Abstract (English)
The non-convex nature of trained neural networks has created significant obstacles in their incorporation into optimization models. In this context, Anderson et al. (2020) provided a framework to obtain the convex hull of the graph of a piecewise linear convex activation function composed with an affine function; this effectively convexifies activations such as the ReLU together with the affine transformation that precedes it. In this article, we contribute to this line of work by developing a recursive formula that yields a tight convexification for the composition of an activation with an affine function for a wide scope of activation functions, namely, convex or ``S-shaped". Our approach can be used to efficiently compute separating hyperplanes or determine that none exists in various settings, including non-polyhedral cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。