arXiv:2412.14075cs.LG2024-12

利用转移原型在线学习,提升不确定环境下的决策鲁棒性。

Online MDP with Transition Prototypes: A Robust Adaptive Approach

  • 基于有限转移原型构建可自适应更新的不确定性集
  • 实现子线性遗憾,且在数据有限时表现更优
  • 适合需快速响应的实时决策场景

本文研究一种在线鲁棒马尔可夫决策过程(MDP),已知底层转移核的有限个原型。通过自适应更新原型的模糊集,提出一种算法,在高效识别真实转移核的同时保证相应鲁棒策略的性能。具体地,我们证明了后续最优鲁棒策略具有子线性后悔率,并提供早期停止机制及价值函数最坏情况性能界。数值实验表明,该方法在数据有限的早期阶段显著优于现有方法。本工作通过引入转移概率的先验信息与在线学习结合,为不确定性下的决策提供了理论支持与实用算法。

原文摘要 · Abstract (English)

In this work, we consider an online robust Markov Decision Process (MDP) where we have the information of finitely many prototypes of the underlying transition kernel. We consider an adaptively updated ambiguity set of the prototypes and propose an algorithm that efficiently identifies the true underlying transition kernel while guaranteeing the performance of the corresponding robust policy. To be more specific, we provide a sublinear regret of the subsequent optimal robust policy. We also provide an early stopping mechanism and a worst-case performance bound of the value function. In numerical experiments, we demonstrate that our method outperforms existing approaches, particularly in the early stage with limited data. This work contributes to robust MDPs by considering possible prior information about the underlying transition probability and online learning, offering both theoretical insights and practical algorithms for improved decision-making under uncertainty.

强化学习鲁棒决策在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。