arXiv:2412.05783cs.LGstat.ML2024-12NeurIPS被引 9

提出双向去混淆方法,解决因果强化学习中未观测混杂因素的策略评估问题。

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

  • 基于面板数据中的双向固定效应思想,构建双向未观测混杂假设。
  • 设计神经张量网络联合学习未观测混杂因子与系统动态。
  • 适用于存在隐藏变量的强化学习策略评估,尤其适合工业级决策场景。

本文研究在存在未测量混杂因素情况下的离策略评估(OPE)问题。受面板数据文献中广泛使用的双向固定效应回归模型启发,我们提出了一个双向未测量混杂假设,用于建模因果强化学习中的系统动态,并开发了双向去混淆算法。该算法通过神经张量网络同时学习未测量混杂因子和系统动态,进而构建基于模型的估计器,实现一致的策略价值估计。我们通过理论分析和数值实验验证了所提估计器的有效性。

原文摘要 · Abstract (English)

This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured confounding assumption to model the system dynamics in causal reinforcement learning and develop a two-way deconfounder algorithm that devises a neural tensor network to simultaneously learn both the unmeasured confounders and the system dynamics, based on which a model-based estimator can be constructed for consistent policy value estimation. We illustrate the effectiveness of the proposed estimator through theoretical results and numerical experiments.

因果强化学习离策略评估未观测混杂

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。