提出在复杂观测下高效学习潜在动态的理论框架与算法。
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
- 定义潜在传播可覆盖性,确保统计可学习性
- 设计两类可观测到潜在的可证明高效还原方法
- 适合研究强化学习理论与复杂环境建模的学者
现实世界中的强化学习常面临高维观测但潜在动态简单的场景。然而,在非小规模潜在空间下,相关统计要求与算法原理仍不明确。本文从统计与算法角度研究一般潜在动态下的强化学习问题。统计上,主要负结果表明:大多数已有函数逼近设置在结合丰富观测时变得不可行;正结果则提出‘潜在传播可覆盖性’作为保证统计可处理性的通用条件。算法上,本文构建了两类可证明高效的可观测到潜在还原方法——一类依赖事后潜动态观测(LADZ23),另一类基于自预测潜模型估计(SAGHCB20)。整体成果为潜在动态下强化学习的统一理论奠定基础。
原文摘要 · Abstract (English)
Real-world applications of reinforcement learning often involve environments where agents operate on complex, high-dimensional observations, but the underlying (''latent'') dynamics are comparatively simple. However, outside of restrictive settings such as small latent spaces, the fundamental statistical requirements and algorithmic principles for reinforcement learning under latent dynamics are poorly understood. This paper addresses the question of reinforcement learning under $\textit{general}$ latent dynamics from a statistical and algorithmic perspective. On the statistical side, our main negative result shows that most well-studied settings for reinforcement learning with function approximation become intractable when composed with rich observations; we complement this with a positive result, identifying latent pushforward coverability as a general condition that enables statistical tractability. Algorithmically, we develop provably efficient observable-to-latent reductions -- that is, reductions that transform an arbitrary algorithm for the latent MDP into an algorithm that can operate on rich observations -- in two settings: one where the agent has access to hindsight observations of the latent dynamics [LADZ23], and one where the agent can estimate self-predictive latent models [SAGHCB20]. Together, our results serve as a first step toward a unified statistical and algorithmic theory for reinforcement learning under latent dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。