提出半参数双强化学习,提升长期因果推断的稳定性与效率。
Semiparametric Double Reinforcement Learning with Applications to Long-Term Causal Inference
- 在无限时域Q函数上施加半参数约束,提升估计稳定性。
- 正确设定下可达到半参数效率边界,优于传统非参数方法。
- 适用于长期干预效果评估,尤其适合随机实验后的长期推断。
双强化学习(DRL)可在非参数马尔可夫决策过程(MDP)中实现策略价值的高效离策略推断,但当时间重叠较弱且占用比率维度较高时,完全非参数估计器可能不稳定。这一限制在随机实验的长期因果推断中尤为显著:虽然随机化保证了治疗分配的重叠,但无法保证持续干预后未来状态轨迹的重叠。本文针对无限时域Q函数的连续线性泛函,提出半参数双强化学习方法。不假设奖励与转移律为线性MDP结构,而是对解折扣贝尔曼方程的Q函数本身施加工作半参数限制。若设定正确,该方法可比无约束DRL更高效,同时允许丰富、可能无限维的模型。为避免依赖正确设定,通过加权贝尔曼残差最小化定义估计目标,其投影目标在误设下仍具意义,并在正确设定下恢复原泛函。我们推导了有效影响函数与效率界,构建了模型鲁棒的自动去偏估计量,并发展了最小极大准则用于估计Q函数和瑞兹函数。在正确设定下,最优加权版本可达受限模型中的半参数效率边界。
原文摘要 · Abstract (English)
Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs), but fully nonparametric estimators can be unstable when intertemporal overlap is weak and occupancy ratios are high-dimensional. This limitation is especially relevant for long-term causal inference from randomized experiments: randomization ensures overlap in treatment assignment, but not over future state trajectories induced by continued intervention use. We develop semiparametric DRL for continuous linear functionals of the infinite-horizon $Q$-function. Rather than impose linear MDP structure on the reward and transition laws, we place working semiparametric restrictions on the $Q$-function itself, the solution of the discounted Bellman equation. When correct, these restrictions can improve efficiency relative to unrestricted DRL while allowing rich, possibly infinite-dimensional models. To avoid relying on correct specification, we define the estimand through weighted Bellman-residual minimization. The resulting projection target remains meaningful under misspecification and recovers the original functional under correct specification. For this class of parameters, we derive efficient influence functions and efficiency bounds, construct model-robust automatically debiased estimators, and develop minimax criteria for estimating the $Q$- and Riesz functions. Under correct specification, optimally weighted versions attain the semiparametric efficiency bound in the restricted model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。