同时优化多个平均收益并保证局部稳定性的首个高效算法
Multiple Mean-Payoff Optimization under Local Stability Constraints
- 提出基于窗口滑动的局部稳定性约束机制
- 首次实现马尔可夫决策过程下的多目标平均收益优化
- 适用于需要稳定性能的实时控制系统设计
长期每步平均回报(均值回报)是描述离散系统性能与可靠性的重要工具。在随机与博弈模型中,同时优化多个均值回报的控制器构造问题已被深入研究。然而,现有控制器普遍存在均值回报不稳定的缺陷——即在运行过程中,通过有限滑动窗口计算的每步平均奖励存在显著波动。由于在局部稳定性约束下同时优化多个均值回报的问题具有计算困难性,现有工作即使对非随机模型如双人博弈也未能提供可实际应用的算法。本文首次设计并评估了适用于马尔可夫决策过程的高效可扩展解决方案。
原文摘要 · Abstract (English)
The long-run average payoff per transition (mean payoff) is the main tool for specifying the performance and dependability properties of discrete systems. The problem of constructing a controller (strategy) simultaneously optimizing several mean payoffs has been deeply studied for stochastic and game-theoretic models. One common issue of the constructed controllers is the instability of the mean payoffs, measured by the deviations of the average rewards per transition computed in a finite "window" sliding along a run. Unfortunately, the problem of simultaneously optimizing the mean payoffs under local stability constraints is computationally hard, and the existing works do not provide a practically usable algorithm even for non-stochastic models such as two-player games. In this paper, we design and evaluate the first efficient and scalable solution to this problem applicable to Markov decision processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。