将马尔可夫决策过程融入多阶段随机规划,提升建模灵活性与求解效率。
MDP modeling for multi-stage stochastic programs
- 用扩展的策略图建模决策依赖的不确定性
- 提出新变体的随机对偶动态规划处理非凸性
- 适用于连续状态动作空间的复杂决策问题
我们研究一类多阶段随机规划,其融合了马尔可夫决策过程(MDP)的建模特性。该类模型包含具有连续状态和动作空间的结构化MDP。我们扩展了策略图,以支持一步转移概率的决策依赖不确定性,以及有限形式的统计学习。重点在于展示该建模方法的表达能力,通过一系列复杂度递增的实例加以说明。作为求解方法,我们开发了随机对偶动态规划的新变体,包括用于处理非凸性的近似方法。
原文摘要 · Abstract (English)
We study a class of multi-stage stochastic programs, which incorporate modeling features from Markov decision processes (MDPs). This class includes structured MDPs with continuous action and state spaces. We extend policy graphs to include decision-dependent uncertainty for one-step transition probabilities as well as a limited form of statistical learning. We focus on the expressiveness of our modeling approach, illustrating ideas with a series of examples of increasing complexity. As a solution method, we develop new variants of stochastic dual dynamic programming, including approximations to handle non-convexities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。