arXiv:2509.14203math.OCcs.LG2025-09被引 5

提出常增益平均奖励鲁棒MDP的贝尔曼最优性理论,填补长期决策控制空白

Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain

  • 基于常增益鲁棒贝尔曼方程分析平均奖励最优性
  • 给出解存在性条件及与最优策略的关系
  • 适用于存在信息不对称的长期运营决策场景

在鲁棒马尔可夫决策过程(MDP)的学习与最优控制中,现有理论、算法和应用多聚焦于有限时域或折扣模型。而长期平均奖励形式虽在运筹学与管理领域自然常见,却仍研究不足,主要因动态规划基础技术复杂且理解不充分,诸多基本问题尚未解决。本文通过分析常增益设定,迈向平均奖励鲁棒MDP的通用框架。研究控制器与S-矩形对手间可能存在信息不对称的平均奖励鲁棒控制问题。核心围绕常增益鲁棒贝尔曼方程,探讨解的存在性及其与最优平均奖励的关系。具体识别出何时鲁棒贝尔曼方程的解能表征最优平均奖励与平稳策略,并提供单侧弱通信条件以保证解的存在。这些发现扩展了平均奖励鲁棒MDP的动态规划理论,为操作环境中基于长期平均准则的鲁棒动态决策奠定了基础。

原文摘要 · Abstract (English)

Learning and optimal control under robust Markov decision processes (MDPs) have received increasing attention, yet most existing theory, algorithms, and applications focus on finite-horizon or discounted models. Long-run average-reward formulations, while natural in many operations research and management contexts, remain underexplored. This is primarily because the dynamic programming foundations are technically challenging and only partially understood, with several fundamental questions remaining open. This paper steps toward a general framework for average-reward robust MDPs by analyzing the constant-gain setting. We study the average-reward robust control problem with possible information asymmetries between the controller and an S-rectangular adversary. Our analysis centers on the constant-gain robust Bellman equation, examining both the existence of solutions and their relationship to the optimal average reward. Specifically, we identify when solutions to the robust Bellman equation characterize the optimal average reward and stationary policies, and we provide one-sided weak communication conditions ensuring solutions' existence. These findings expand the dynamic programming theory for average-reward robust MDPs and lay a foundation for robust dynamic decision making under long-run average criteria in operational environments.

强化学习鲁棒控制动态规划平均奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。