arXiv:2603.02862cs.LG2026-03被引 1

针对外生状态的强化学习,显著提升样本效率。

Learning in Markov Decision Processes with Exogenous Dynamics

  • 区分可控与外生状态,利用结构化假设改进学习。
  • 后悔值上界仅依赖外生状态空间大小,性能大幅提升。
  • 适合状态复杂但部分可控的现实场景,如机器人控制。

强化学习算法通常针对通用马尔可夫决策过程(MDP),其中任意状态-动作对均可导致任意转移分布。但在许多实际系统中,仅有部分状态变量受智能体动作直接影响,其余成分由外生动态演变,且主导了大部分随机性。本文研究一类具有外生状态组件的结构化MDP,其转移独立于智能体动作。我们证明,利用该结构可获得显著更优的学习保证,后悔值上界的主导项仅与外生状态空间大小相关。进一步建立了匹配的下界,表明该依赖关系在信息论上最优。最后,在经典玩具环境和真实世界启发的场景中进行实证验证,结果表明相比标准强化学习方法,本方法在样本效率上实现显著提升。

原文摘要 · Abstract (English)

Reinforcement learning algorithms are typically designed for generic Markov Decision Processes (MDPs), where any state-action pair can lead to an arbitrary transition distribution. In many practical systems, however, only a subset of the state variables is directly influenced by the agent's actions, while the remaining components evolve according to exogenous dynamics and account for most of the stochasticity. In this work, we study a structured class of MDPs characterized by exogenous state components whose transitions are independent of the agent's actions. We show that exploiting this structure yields significantly improved learning guarantees, with only the size of the exogenous state space appearing in the leading terms of the regret bounds. We further establish a matching lower bound, showing that this dependence is information-theoretically optimal. Finally, we empirically validate our approach across classical toy settings and real-world-inspired environments, demonstrating substantial gains in sample efficiency compared to standard reinforcement learning methods.

强化学习状态空间样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。