提出首个单次轨迹评估下求解广义效用MDP的方法。
Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning
- 将单次轨迹优化转化为等价MDP,明确最优策略类
- 采用蒙特卡洛树搜索在线规划,实现高效求解
- 实验显示性能优于现有基线,适用于高风险决策
本文首次提出在单次轨迹评估框架下求解无限时域折扣广义效用马尔可夫决策过程(GUMDP)的方法。首先,我们给出了单次轨迹环境下策略优化的基本结论:明确了达到最优所需策略的类别,将原问题转化为一个等价的MDP,并研究了该环境下策略优化的计算难度。其次,我们展示了如何利用在线规划技术,特别是蒙特卡洛树搜索算法,来求解GUMDP。最后,通过实验验证了所提方法在多个基准任务上的优越性能,显著优于现有基线方法。
原文摘要 · Abstract (English)
In this work, we contribute the first approach to solve infinite-horizon discounted general-utility Markov decision processes (GUMDPs) in the single-trial regime, i.e., when the agent's performance is evaluated based on a single trajectory. First, we provide some fundamental results regarding policy optimization in the single-trial regime, investigating which class of policies suffices for optimality, casting our problem as a particular MDP that is equivalent to our original problem, as well as studying the computational hardness of policy optimization in the single-trial regime. Second, we show how we can leverage online planning techniques, in particular a Monte-Carlo tree search algorithm, to solve GUMDPs in the single-trial regime. Third, we provide experimental results showcasing the superior performance of our approach in comparison to relevant baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。