在无限时域通用效用MDP中,采样轨迹数显著影响策略评估结果。
The Number of Trials Matters in Infinite-Horizon General-Utility Markov Decision Processes
- 分析采样轨迹数量对策略性能评估的影响机制
- 证明有限与无限轨迹下评估结果存在可量化偏差
- 适用于关注长期策略评估稳定性的强化学习研究者
通用效用马尔可夫决策过程(GUMDP)框架通过考虑策略诱导的状态-动作对访问频率来扩展标准MDP框架。本文首次分析了无限时域GUMDP中采样轨迹数(即试验次数)的影响。研究表明,与标准MDP不同,轨迹数量在无限时域GUMDP中起关键作用,给定策略的期望性能通常依赖于采样轨迹数。本文分别研究折扣型和平均型GUMDP,其中目标函数分别依赖于折扣访问频率和平均访问频率。首先,在折扣型GUMDP下,推导出有限与无限轨迹设定间误差的上下界;其次,在平均型GUMDP中,分析不同类别问题对有限与无限轨迹评估差异的影响;最后,通过一系列实验验证结论,揭示轨迹数量与底层GUMDP结构如何共同影响策略评估效果。
原文摘要 · Abstract (English)
The general-utility Markov decision processes (GUMDPs) framework generalizes the MDPs framework by considering objective functions that depend on the frequency of visitation of state-action pairs induced by a given policy. In this work, we contribute with the first analysis on the impact of the number of trials, i.e., the number of randomly sampled trajectories, in infinite-horizon GUMDPs. We show that, as opposed to standard MDPs, the number of trials plays a key-role in infinite-horizon GUMDPs and the expected performance of a given policy depends, in general, on the number of trials. We consider both discounted and average GUMDPs, where the objective function depends, respectively, on discounted and average frequencies of visitation of state-action pairs. First, we study policy evaluation under discounted GUMDPs, proving lower and upper bounds on the mismatch between the finite and infinite trials formulations for GUMDPs. Second, we address average GUMDPs, studying how different classes of GUMDPs impact the mismatch between the finite and infinite trials formulations. Third, we provide a set of empirical results to support our claims, highlighting how the number of trajectories and the structure of the underlying GUMDP influence policy evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。