arXiv:2505.00304stat.MLcs.LG2025-05被引 5

解决连续动作强化学习中未观测混杂因素的策略评估难题

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding

  • 基于新识别结果实现无限时域下策略价值的非参数估计
  • 提出极小极大估计器和策略梯度算法,找到最优策略
  • 理论证明一致性和误差界,适合真实数据应用

本文针对存在未观测混杂因素的连续动作强化学习中的离线策略学习问题。现有研究多集中于部分可观测马尔可夫决策过程(POMDP)下的策略评估,并假设动作空间为离散。本文在无限时域框架下建立了新的识别结果,实现对目标策略策略价值的非参数估计。基于该识别结果,提出极小极大估计器,并设计基于策略梯度的算法以识别使估计策略价值最大化的类内最优策略。进一步提供了所获最优策略的一致性、有限样本误差界及遗憾界等理论结果。通过大量仿真与基于德国家庭面板数据的真实世界应用,验证了方法的有效性。

原文摘要 · Abstract (English)

This paper addresses the challenge of offline policy learning in reinforcement learning with continuous action spaces when unmeasured confounders are present. While most existing research focuses on policy evaluation within partially observable Markov decision processes (POMDPs) and assumes discrete action spaces, we advance this field by establishing a novel identification result to enable the nonparametric estimation of policy value for a given target policy under an infinite-horizon framework. Leveraging this identification, we develop a minimax estimator and introduce a policy-gradient-based algorithm to identify the in-class optimal policy that maximizes the estimated policy value. Furthermore, we provide theoretical results regarding the consistency, finite-sample error bound, and regret bound of the resulting optimal policy. Extensive simulations and a real-world application using the German Family Panel data demonstrate the effectiveness of our proposed methodology.

强化学习连续动作混杂因素离线策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。