arXiv:2603.06946cs.LGmath.OC2026-03被引 1

提出联合马尔可夫决策过程,统一建模多动作间的因果依赖关系。

Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments

  • 引入多动作生成接口,显式建模不同动作在相同随机条件下的结果关联
  • 构建支持瞬时回报矩的贝尔曼算子,实现动态规划与增量算法收敛
  • 适用于需分析动作间依赖的强化学习场景,如多臂赌博机、策略比较

强化学习中的许多分布量(如差距分布、优势概率)本质上是跨动作联合的。然而经典马尔可夫决策过程(MDP)仅定义边际分布,未指定同一状态下多个动作的反事实一步结果之间的联合分布。本文研究具有耦合动力学的环境,采用多动作生成接口,在共享外生随机性下采样多个动作的反事实一步结果。提出联合MDP(JMDP)形式化框架,通过扩展标准MDP并引入多动作样本转移模型,明确指定一步反事实结果的耦合结构,同时保持原有MDP交互作为边际观测。采用一步耦合范式,将动作间依赖限制在查询状态的即时反事实结果上。在此设定下,推导了第n阶回报矩的贝尔曼算子,为动态规划与增量算法提供收敛保证。

原文摘要 · Abstract (English)

Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority. However, the classical Markov decision process (MDP) formalism specifies only marginal laws and leaves the joint law of counterfactual one-step outcomes across multiple possible actions at a state unspecified. We study coupled-dynamics environments with a multi-action generative interface which can sample counterfactual one-step outcomes for multiple actions under shared exogenous randomness. We propose joint MDPs (JMDPs) as a formalism for such environments by augmenting an MDP with a multi-action sample transition model which specifies a coupling of one-step counterfactual outcomes, while preserving standard MDP interaction as marginal observations. We adopt and formalize a one-step coupling regime where dependence across actions is confined to immediate counterfactual outcomes at the queried state. In this regime, we derive Bellman operators for $n$th-order return moments, providing dynamic programming and incremental algorithms with convergence guarantees.

强化学习联合分布动态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。