arXiv:2512.20582cs.LGcs.GT2025-12

将ReLU神经网络转化为零和回合制博弈,揭示其输出的博弈论本质。

Relu and softplus neural nets as zero-sum turn-based games

  • 用反向博弈视角解释网络输出,输入为终局奖励
  • 推导出基于路径积分的离散费曼-卡茨公式
  • 可用于验证鲁棒性或反向设计网络参数

我们证明,ReLU神经网络的输出可被解释为一个零和、回合制、停止博弈的值,称为ReLU网络博弈。该博弈运行方向与网络相反,网络输入作为博弈的终局奖励。实际上,评估网络等价于对博弈值执行谢普利-贝尔曼逆向递推。利用博弈值表示为路径测度下期望总收益的形式,我们推导出网络输出的离散费曼-卡茨型路径积分公式。该博弈论表达式可用于从输入边界推导输出边界,利用谢普利算子的单调性;也可通过策略作为证书来验证鲁棒性。此外,训练神经网络变为逆博弈问题:给定终局奖励及其对应值,求解能重现这些结果的博弈转移概率和奖励。最后,我们证明类似方法也适用于含Softplus激活函数的网络,其中ReLU网络博弈被其熵正则化版本替代。

原文摘要 · Abstract (English)

We show that the output of a ReLU neural network can be interpreted as the value of a zero-sum, turn-based, stopping game, which we call the ReLU net game. The game runs in the direction opposite to that of the network, and the input of the network serves as the terminal reward of the game. In fact, evaluating the network is the same as running the Shapley-Bellman backward recursion for the value of the game. Using the expression of the value of the game as an expected total payoff with respect to the path measure induced by the transition probabilities and a pair of optimal policies, we derive a discrete Feynman-Kac-type path-integral formula for the network output. This game-theoretic representation can be used to derive bounds on the output from bounds on the input, leveraging the monotonicity of Shapley operators, and to verify robustness properties using policies as certificates. Moreover, training the neural network becomes an inverse game problem: given pairs of terminal rewards and corresponding values, one seeks transition probabilities and rewards of a game that reproduces them. Finally, we show that a similar approach applies to neural networks with Softplus activation functions, where the ReLU net game is replaced by its entropic regularization.

博弈论神经网络鲁棒性路径积分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。