用扩散模型学习部分可观测环境中的因果状态,提升强化学习在干扰下的表现。
Learning Causal States Under Partial Observability and Perturbation
- 引入异步扩散模型与新双模拟度量,揭示部分可观测环境的因果结构。
- 实验显示在Roboschool任务上收益提升至少14.18%,优于基线方法。
- 首个结合理论保障与实际应用的因果状态扩散框架,适合复杂干扰场景研究者。
强化学习在不完整且噪声干扰的观测下做出决策面临严峻挑战,尤其在受扰动的部分可观测马尔可夫决策过程(P²OMDP)中。现有方法难以同时应对观测不全与扰动问题。我们提出 extit{CaDiff} 框架,通过引入新型异步扩散模型(ADM)与双模拟度量,增强任意强化学习算法对 P²OMDP 隐含因果结构的识别能力。ADM 支持正向与反向过程步数不同,将环境扰动视为可通过扩散过程抑制的噪声。双模拟度量量化部分可观测环境与其因果对应物间的相似性。理论上,我们推导出价值函数近似误差的上界,体现奖励与转移模型近似误差间的合理权衡。在 Roboschool 任务上的实验表明,相比基线方法,CaDiff 至少提升 14.18% 的回报。该框架是首个兼具理论严谨性与实用性的基于扩散模型的因果状态近似方法。
原文摘要 · Abstract (English)
A critical challenge for reinforcement learning (RL) is making decisions based on incomplete and noisy observations, especially in perturbed and partially observable Markov decision processes (P$^2$OMDPs). Existing methods fail to mitigate perturbations while addressing partial observability. We propose \textit{Causal State Representation under Asynchronous Diffusion Model (CaDiff)}, a framework that enhances any RL algorithm by uncovering the underlying causal structure of P$^2$OMDPs. This is achieved by incorporating a novel asynchronous diffusion model (ADM) and a new bisimulation metric. ADM enables forward and reverse processes with different numbers of steps, thus interpreting the perturbation of P$^2$OMDP as part of the noise suppressed through diffusion. The bisimulation metric quantifies the similarity between partially observable environments and their causal counterparts. Moreover, we establish the theoretical guarantee of CaDiff by deriving an upper bound for the value function approximation errors between perturbed observations and denoised causal states, reflecting a principled trade-off between approximation errors of reward and transition-model. Experiments on Roboschool tasks show that CaDiff enhances returns by at least 14.18\% compared to baselines. CaDiff is the first framework that approximates causal states using diffusion models with both theoretical rigor and practicality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。