提出新方法建模多动作间的依赖关系,提升复杂环境下的决策性能。
Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

- 基于联合马尔可夫决策过程建模多动作间依赖关系
- 证明在特定条件下算法迭代收敛到最优联合回报分布
- 适用于需考虑多种动作后果关联的强化学习场景
耦合动态环境揭示了在共同外生随机实现下,多个反事实动作可能带来的单步结果。传统马尔可夫决策过程仅能分析每个动作的边缘分布,忽略了这些反事实结果之间的依赖关系。联合马尔可夫决策过程(JMDP)形式化保留了这种依赖。先前工作建立了该形式化框架并解决了固定策略下的联合矩评估问题。本文发展了最优控制方法:定义了一个非参数化的分布贝尔曼最优算子,并证明当诱导的边际MDP具有唯一最优策略时,其迭代在Wasserstein距离下收敛至最优联合回报分布。对于前两阶矩,我们建立了在更弱条件下仍可收敛的保证,允许存在多个均值最优动作,只要它们的冲突解决方式共享同一个二阶矩不动点。此外,还推导了用于神经网络近似的采样目标。
原文摘要 · Abstract (English)
Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism allows one to reason about the marginal law of each action but discards dependence across these counterfactual outcomes. The Joint Markov decision process (JMDP) formalism preserves that dependence. Prior work established the formalism and solved the fixed-policy joint moment evaluation problem in JMDPs. This paper develops optimal-control methods. We define a nonparametric distributional Bellman optimality operator for JMDPs, and prove that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law. For the first two moments, we establish convergence under a weaker condition that permits several mean-optimal actions as long as their tie resolutions share a second-moment fixed point. We also derive sampled targets for neural approximation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。