arXiv:2512.07528cs.LGcs.AI2025-12被引 1

解决观测不到上下文时模型强化学习的偏差问题

Model-Based Reinforcement Learning Under Confounding

  • 利用可观察的代理变量识别混淆下的真实奖励期望
  • 构建一致的替代马尔可夫决策过程,支持可靠规划
  • 适合在上下文不可获取的现实场景中使用

我们研究在上下文未被观测且导致离线数据混淆的上下文马尔可夫决策过程(C-MDP)中的基于模型强化学习。传统模型学习方法在此类场景下根本上不一致,因行为策略下的转移与奖励机制不对应于评估状态策略所需的干预量。为此,我们采用一种近似离策略评估方法,在代理变量满足轻微可逆性条件下,仅利用可观测的状态-动作-奖励轨迹识别混淆的奖励期望。结合行为平均转移模型,该构造生成一个贝尔曼算子明确定义且对状态策略一致的替代MDP,可无缝融入最大因果熵(MaxCausalEnt)模型学习框架。所提方法实现了在上下文无法观测、不可得或难以采集的环境中,有原则的模型学习与规划。

原文摘要 · Abstract (English)

We investigate model-based reinforcement learning in contextual Markov decision processes (C-MDPs) in which the context is unobserved and induces confounding in the offline dataset. In such settings, conventional model-learning methods are fundamentally inconsistent, as the transition and reward mechanisms generated under a behavioral policy do not correspond to the interventional quantities required for evaluating a state-based policy. To address this issue, we adapt a proximal off-policy evaluation approach that identifies the confounded reward expectation using only observable state-action-reward trajectories under mild invertibility conditions on proxy variables. When combined with a behavior-averaged transition model, this construction yields a surrogate MDP whose Bellman operator is well defined and consistent for state-based policies, and which integrates seamlessly with the maximum causal entropy (MaxCausalEnt) model-learning framework. The proposed formulation enables principled model learning and planning in confounded environments where contextual information is unobserved, unavailable, or impractical to collect.

强化学习因果推断模型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。