通过最优响应流求解非凸概率分布优化,实现全局收敛。
Non-convex entropic mean-field optimization via Best Response flow
- 用最优响应流处理带熵正则化的非凸问题,设计可收缩的正则项。
- 在特定正则化下,保证唯一全局最优解存在且可通过流迭代逼近。
- 适用于马尔可夫决策与博弈中的策略优化,尤其适合软最大参数化。
研究在概率测度空间上最小化非凸泛函的问题,该问题通过相对于固定参考测度的相对熵(KL散度)进行正则化,以及对应的熵正则化非凸-非凹极小极大问题。我们采用最优响应流(也称虚构博弈流),分析其收敛性如何受目标泛函非凸程度、正则化参数及参考测度尾部行为的影响。特别地,我们展示了如何根据非凸泛函选择合适的正则项,使得最优响应算子在$L^1$-Wasserstein距离下成为压缩映射,从而确保其唯一不动点存在,并被证明是原优化问题的唯一全局最小值。这一结果扩展了近期关于最优响应流应用于任意参考测度和任意正则化参数下的凸优化的研究。我们的结果精确说明了如何在牺牲部分一般性的情况下放松凸性假设。此外,我们展示了这些结果在强化学习中应用于马尔可夫决策过程和马尔可夫博弈的策略优化,特别是在平均场框架下使用软最大参数化的情形。
原文摘要 · Abstract (English)
We study the problem of minimizing non-convex functionals on the space of probability measures, regularized by the relative entropy (KL divergence) with respect to a fixed reference measure, as well as the corresponding problem of solving entropy-regularized non-convex-non-concave min-max problems. We utilize the Best Response flow (also known in the literature as the fictitious play flow) and study how its convergence is influenced by the relation between the degree of non-convexity of the functional under consideration, the regularization parameter and the tail behaviour of the reference measure. In particular, we demonstrate how to choose the regularizer, given the non-convex functional, so that the Best Response operator becomes a contraction with respect to the $L^1$-Wasserstein distance, which ensures the existence of its unique fixed point that is then shown to be the unique global minimizer for our optimization problem. This extends recent results where the Best Response flow was applied to solve convex optimization problems regularized by the relative entropy with respect to arbitrary reference measures, and with arbitrary values of the regularization parameter. Our results explain precisely how the assumption of convexity can be relaxed, at the expense of making a specific choice of the regularizer. Additionally, we demonstrate how these results can be applied in reinforcement learning in the context of policy optimization for Markov Decision Processes and Markov games with softmax parametrized policies in the mean-field regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。