arXiv:2502.03016math.OCcs.LG2025-02被引 5

研究ReLU神经网络优化难题,发现模型冗余会显著增加求解成本。

An analysis of optimization problems involving ReLU neural networks

  • 通过剪裁、正则化等方法改进神经网络训练,降低优化复杂度。
  • 实测显示模型层数增加导致大M系数指数级增长,计算耗时大幅上升。
  • 适合关注神经网络可优化性与实际部署效率的研究者阅读。

求解嵌入ReLU激活函数神经网络的混合整数优化问题极具挑战性。与这些函数相关的二元变量松弛所产生之Big-M系数随网络层数呈指数级增长。本文综述并提出多种分析与改善混合整数规划求解器性能的方法,包括训练阶段的剪裁变体与正则化技术,以及基于优化的边界紧缩和针对给定ReLU网络的新缩放策略。我们在文献中的三个基准问题上对这些方法进行数值比较,以线性区域数量、稳定神经元比例及整体计算开销为评估指标。主要发现是:模型冗余虽常被期望,但会显著增加相关优化问题的计算成本,且该权衡关系得到量化验证。

原文摘要 · Abstract (English)

Solving mixed-integer optimization problems with embedded neural networks with ReLU activation functions is challenging. Big-M coefficients that arise in relaxing binary decisions related to these functions grow exponentially with the number of layers. We survey and propose different approaches to analyze and improve the run time behavior of mixed-integer programming solvers in this context. Among them are clipped variants and regularization techniques applied during training as well as optimization-based bound tightening and a novel scaling for given ReLU networks. We numerically compare these approaches for three benchmark problems from the literature. We use the number of linear regions, the percentage of stable neurons, and overall computational effort as indicators. As a major takeaway we observe and quantify a trade-off between the often desired redundancy of neural network models versus the computational costs for solving related optimization problems.

神经网络优化混合整数规划ReLU计算效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。