提出新方法解决非凸非凹极小极大问题,用于分布鲁棒优化与对抗训练。
A stochastic smoothing framework for nonconvex-nonconcave minEmax problems with applications to Wasserstein distributionally robust optimization
- 用对数均值指数平滑处理价值函数,构建随机近端梯度算法。
- 理论证明收敛至Clarke平稳点,数值实验显示性能优于现有基线。
- 适合做分布鲁棒优化、对抗训练等复杂场景的优化求解。
我们研究一类随机非光滑优化问题,其中外变量最小化一个逐点最大值的期望。此类最小化-期望-最大化(minEmax)问题出现在Wasserstein分布鲁棒优化和对抗鲁棒训练中,当底层分布非经验分布时,通常无法重构成有限维极小极大问题。本文基于对数均值指数平滑方法,提出一种随机平滑近端梯度算法。在紧性与Lipschitz型假设下,我们给出了关于Goldstein平稳性的非渐近分析,并证明该方法生成的几乎必然聚类点是原始问题的Clarke平稳点;由Clarke正则性,该点也是方向平稳点。在报童问题、鲁棒回归及对抗鲁棒学习问题上的数值实验表明,所提方法在性能上可与现有基线相媲美。
原文摘要 · Abstract (English)
We study a class of stochastic nonsmooth optimization problems in which an outer variable minimizes the expectation of a pointwise maximum. This minimization--expectation--maximization (minEmax) problem arises in Wasserstein distributionally robust optimization and adversarially robust training, and it cannot in general be reformulated as a finite-dimensional minimax problem when the underlying distribution is not empirical. We propose a stochastic smoothing proximal gradient method based on log-mean-exp smoothing of the value function. Under compactness and Lipschitz-type assumptions, we present nonasymptotic analysis in terms of Goldstein stationarity and show that every almost-sure cluster point generated by our method is a Clarke stationary point; by Clarke regularity, such a point is also directional stationary for the original problem. Numerical experiments on newsvendor, robust regression, and adversarially robust learning problems show that the proposed method is competitive with existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。