用因果模型提升环境变化下的智能体规划能力
Planning under Distribution Shifts with Causal POMDPs
- 基于因果POMDP构建可应对分布偏移的规划框架
- 保持价值函数分段线性凸性,保障规划可计算性
- 适合需要适应动态环境的机器人与强化学习应用
真实世界中,规划常受分布偏移挑战。同一环境下训练的模型在状态分布或环境动态变化后可能失效,导致策略失败。本文提出一种基于因果知识的POMDP框架,将环境变化建模为对因果POMDP的干预。该框架支持评估假设变化下的计划,并主动识别被改变的环境组件。我们展示了如何维护并更新关于隐状态和底层领域双重信念,并证明价值函数在扩展后的信念空间中仍保持分段线性凸(PWLC)性质。这一特性确保了使用α向量法进行规划的可计算性,在分布偏移下依然可行。
原文摘要 · Abstract (English)
In the real world, planning is often challenged by distribution shifts. As such, a model of the environment obtained under one set of conditions may no longer remain valid as the distribution of states or the environment dynamics change, which in turn causes previously learned strategies to fail. In this work, we propose a theoretical framework for planning under partial observability using Partially Observable Markov Decision Processes (POMDPs) formulated using causal knowledge. By representing shifts in the environment as interventions on this causal POMDP, the framework enables evaluating plans under hypothesized changes and actively identifying which components of the environment have been altered. We show how to maintain and update a belief over both the latent state and the underlying domain, and we prove that the value function remains piecewise linear and convex (PWLC) in this augmented belief space. Preservation of PWLC under distribution shifts has the advantage of maintaining the tractability of planning via $α$-vector-based POMDP methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。