提出交叉拟合方法,提升隐变量强化学习中模型估计的准确性与数据效率。
Cross-fitted Proximal Learning for Model-Based Reinforcement Learning

- 将桥接函数估计建模为带干扰项的条件矩约束问题
- 通过K折交叉拟合提升数据利用效率,降低估计偏差
- 适用于存在隐藏混淆的不完全可观测强化学习场景
基于模型的强化学习因显式建模奖励与转移函数并支持模拟回溯而备受关注。然而,在存在隐藏混淆的离线设置中,直接从观测数据学习的模型可能存在偏差。这一挑战在部分可观测系统中尤为显著,其中潜在因素可能共同影响动作、奖励和未来观测。近期研究发现,此类混淆的部分可观测马尔可夫决策过程(POMDP)中的策略评估可转化为估计满足条件矩约束(CMR)的奖励-观测发射和观测-转移桥接函数。本文研究这些桥接函数的统计估计。我们将桥接学习形式化为一个包含条件均值嵌入和条件密度作为干扰项的CMR问题,并提出现有两阶段桥接估计器的K折交叉拟合扩展。所提方法保持原有桥接识别策略的同时,比单一样本划分更高效地利用数据。我们还推导了交叉拟合估计器的最优比较界,并将误差分解为由干扰项估计引起的第I阶段项与由经验平均引起的第II阶段项。
原文摘要 · Abstract (English)
Model-based reinforcement learning is attractive for sequential decision-making because it explicitly estimates reward and transition models and then supports planning through simulated rollouts. In offline settings with hidden confounding, however, models learned directly from observational data may be biased. This challenge is especially pronounced in partially observable systems, where latent factors may jointly affect actions, rewards, and future observations. Recent work has shown that policy evaluation in such confounded partially observable Markov decision processes (POMDPs) can be reduced to estimating reward-emission and observation-transition bridge functions satisfying conditional moment restrictions (CMRs). In this paper, we study the statistical estimation of these bridge functions. We formulate bridge learning as a CMR problem with nuisance objects given by a conditional mean embedding and a conditional density. We then develop a $K$-fold cross-fitted extension of the existing two-stage bridge estimator. The proposed procedure preserves the original bridge-based identification strategy while using the available data more efficiently than a single sample split. We also derive an oracle-comparator bound for the cross-fitted estimator and decompose the resulting error into a Stage I term induced by nuisance estimation and a Stage II term induced by empirical averaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。