提出可证明的带特权信息的部分可观测强化学习方法,解决实际训练中的效率与稳定性问题。
Provable Partially Observable Reinforcement Learning with Privileged Information
- 通过形式化专家蒸馏和非对称演员-评论家机制,揭示其理论局限与适用条件。
- 在确定性滤波条件下,实现多项式样本与计算复杂度,显著提升学习效率。
- 适用于仿真环境中的策略优化,尤其适合有特权信息的多智能体系统设计者。
部分可观测状态通常给强化学习(RL)带来重大挑战。实践中,某些特权信息(如模拟器中的状态访问)已被用于训练并取得显著的实证成功。为更好理解特权信息的优势,本文重新审视并分析几种简单且实用的范式。首先,形式化了专家蒸馏(又称教师-学生学习)的实证范式,揭示其在寻找近似最优策略时的缺陷。随后,识别出部分可观测环境中的一种条件——确定性滤波条件,在此条件下,专家蒸馏可实现样本与计算复杂度均为多项式。此外,研究了另一种常用范式——非对称演员-评论家,并聚焦于可观测部分可观察马尔可夫决策过程这一更具挑战性的设定。提出一种信念加权的非对称演员-评论家算法,具备多项式样本与准多项式计算复杂度,其中关键组件是一个新提出的可证明的信念状态学习预言机,能在模型误设下保持滤波稳定性,可能具有独立价值。最后,还研究了带特权信息的部分可观测多智能体强化学习(MARL)的可证明效率。开发出基于集中训练-分散执行的算法,该框架在前述两种范式中均表现出多项式样本与(准)多项式计算复杂度。相较于少数近期相关理论研究,本工作关注的是实际启发的算法范式,避免使用计算上不可行的预言机。
原文摘要 · Abstract (English)
Partial observability of the underlying states generally presents significant challenges for reinforcement learning (RL). In practice, certain \emph{privileged information}, e.g., the access to states from simulators, has been exploited in training and has achieved prominent empirical successes. To better understand the benefits of privileged information, we revisit and examine several simple and practically used paradigms in this setting. Specifically, we first formalize the empirical paradigm of \emph{expert distillation} (also known as \emph{teacher-student} learning), demonstrating its pitfall in finding near-optimal policies. We then identify a condition of the partially observable environment, the \emph{deterministic filter condition}, under which expert distillation achieves sample and computational complexities that are \emph{both} polynomial. Furthermore, we investigate another useful empirical paradigm of \emph{asymmetric actor-critic}, and focus on the more challenging setting of observable partially observable Markov decision processes. We develop a belief-weighted asymmetric actor-critic algorithm with polynomial sample and quasi-polynomial computational complexities, in which one key component is a new provable oracle for learning belief states that preserve \emph{filter stability} under a misspecified model, which may be of independent interest. Finally, we also investigate the provable efficiency of partially observable multi-agent RL (MARL) with privileged information. We develop algorithms featuring \emph{centralized-training-with-decentralized-execution}, a popular framework in empirical MARL, with polynomial sample and (quasi-)polynomial computational complexities in both paradigms above. Compared with a few recent related theoretical studies, our focus is on understanding practically inspired algorithmic paradigms, without computationally intractable oracles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。