用未来信息提升部分可观测环境下的状态表征能力
Dynamical-VAE-based Hindsight to Learn the Causal Dynamics of Factored-POMDPs
- 基于动态变分自编码器,融合过去、当前和多步未来信息
- 在事实化POMDP中更准确还原隐藏状态转移的因果图
- 适合研究因果推理与部分可观测决策问题的研究者
从部分观测中学习环境潜在动态是机器学习的关键挑战。在部分可观测马尔可夫决策过程(POMDP)中,状态表征通常基于历史观测与动作推断。我们证明,融入未来信息对准确捕捉因果动态、提升状态表征至关重要。为此,我们提出一种动态变分自编码器(DVAE),用于从离线轨迹中学习事实化POMDP中的因果马尔可夫动态。该方法采用扩展的回望框架,将过去、当前及多步未来信息整合到因子化POMDP设置中。实验表明,该方法比基于历史和典型回望模型更能有效揭示隐藏状态转移的因果图。
原文摘要 · Abstract (English)
Learning representations of underlying environmental dynamics from partial observations is a critical challenge in machine learning. In the context of Partially Observable Markov Decision Processes (POMDPs), state representations are often inferred from the history of past observations and actions. We demonstrate that incorporating future information is essential to accurately capture causal dynamics and enhance state representations. To address this, we introduce a Dynamical Variational Auto-Encoder (DVAE) designed to learn causal Markovian dynamics from offline trajectories in a POMDP. Our method employs an extended hindsight framework that integrates past, current, and multi-step future information within a factored-POMDP setting. Empirical results reveal that this approach uncovers the causal graph governing hidden state transitions more effectively than history-based and typical hindsight-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。