对比专家蒸馏与标准强化学习,揭示部分可观测环境下的算法权衡。
To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
- 构建扰动块马尔可夫模型,分析隐状态动态随机性对策略学习的影响。
- 发现最优隐状态策略不总是最佳蒸馏目标,取决于系统随机性水平。
- 为利用额外信息提供新指导,适用于模拟运动等复杂部分可观测任务。
部分可观测性是强化学习中的重大挑战,因需学习依赖历史的复杂策略。近期方法通过特权专家蒸馏(privileged expert distillation)取得成功,即在训练时利用模拟器提供的隐状态信息,学习并模仿最优隐状态马尔可夫策略,从而将“如何看”与“如何行动”分离。尽管该方法比无隐状态信息的标准强化学习更高效,但存在已知失败模式。本文通过一个简洁但具有启发性的理论模型——扰动块马尔可夫决策过程(perturbed Block MDP),以及在高难度模拟运动任务上的受控实验,研究了特权专家蒸馏与标准强化学习之间的算法权衡。主要发现:(1) 该权衡在经验上取决于隐状态动态的随机性,与扰动块MDP中近似可解码性与信念收缩的理论预测一致;(2) 最优隐状态策略并非总为最佳蒸馏目标。研究结果为有效利用特权信息提供了新准则,可能提升众多实际部分可观测领域中的策略学习效率。
原文摘要 · Abstract (English)
Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used privileged expert distillation--which leverages availability of latent state information during training (e.g., from a simulator) to learn and imitate the optimal latent, Markovian policy--to disentangle the task of "learning to see" from "learning to act". While expert distillation is more computationally efficient than RL without latent state information, it also has well-documented failure modes. In this paper--through a simple but instructive theoretical model called the perturbed Block MDP, and controlled experiments on challenging simulated locomotion tasks--we investigate the algorithmic trade-off between privileged expert distillation and standard RL without privileged information. Our main findings are: (1) The trade-off empirically hinges on the stochasticity of the latent dynamics, as theoretically predicted by contrasting approximate decodability with belief contraction in the perturbed Block MDP; and (2) The optimal latent policy is not always the best latent policy to distill. Our results suggest new guidelines for effectively exploiting privileged information, potentially advancing the efficiency of policy learning across many practical partially observable domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。