arXiv:2411.11451cs.AImath.OC2024-11被引 32

让强化学习更稳健:用不确定集代替精确概率

Robust Markov Decision Processes: A Place Where AI and Formal Methods Meet

  • 用不确定性集合替代精确转移概率,增强决策鲁棒性
  • 扩展值迭代与策略迭代,可求解不确定环境下的最优策略
  • 适合关注安全与可靠性的人工智能研究者

马尔可夫决策过程(MDPs)是序列决策问题的标准模型,广泛应用于形式化方法和人工智能领域。然而,其关键假设要求转移概率必须精确已知,这在实际中常不成立。鲁棒马尔可夫决策过程(RMDPs)通过将转移概率定义为某个不确定集中的成员,克服了这一限制。本文提供一份温和的综述,涵盖RMDPs的基本概念,包括语义解释及求解方法,如扩展标准的值迭代和策略迭代算法。我们还讨论了RMDPs与其他模型的关系及其在强化学习、抽象技术等场景中的应用,并指出未来研究面临的挑战。

原文摘要 · Abstract (English)

Markov decision processes (MDPs) are a standard model for sequential decision-making problems and are widely used across many scientific areas, including formal methods and artificial intelligence (AI). MDPs do, however, come with the restrictive assumption that the transition probabilities need to be precisely known. Robust MDPs (RMDPs) overcome this assumption by instead defining the transition probabilities to belong to some uncertainty set. We present a gentle survey on RMDPs, providing a tutorial covering their fundamentals. In particular, we discuss RMDP semantics and how to solve them by extending standard MDP methods such as value iteration and policy iteration. We also discuss how RMDPs relate to other models and how they are used in several contexts, including reinforcement learning and abstraction techniques. We conclude with some challenges for future work on RMDPs.

强化学习鲁棒优化形式化方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。