提出新型元学习器DRQ-learner,提升序列决策中个性化结果预测精度。
An Orthogonal Learner for Individualized Outcomes in Markov Decision Processes
- 基于因果推断构建双重稳健的元学习框架
- 理论保证对误差不敏感且逼近最优效率
- 适用于离散与连续状态空间,兼容神经网络等模型
在个性化医疗等序列决策场景中,长期潜在结果预测极具挑战。现有方法通常缺乏强理论保障。本文从因果推断视角重新审视该问题,提出一种新型元学习器DRQ-learner,具备三项关键性质:(1) 双重稳健性(任一辅助模型误设仍可保证有效推断),(2) Neyman正交性(对辅助函数的一阶估计误差不敏感),(3) 接近最优效率(渐近表现如同已知真实辅助函数)。该方法适用于离散与连续状态空间,可与任意机器学习模型(如神经网络)结合使用。数值实验验证了其优于现有最先进基线的表现。
原文摘要 · Abstract (English)
Predicting individualized potential outcomes in sequential decision-making is central for optimizing therapeutic decisions in personalized medicine (e.g., which dosing sequence to give to a cancer patient). However, predicting potential outcomes over long horizons is notoriously difficult. Existing methods that break the curse of the horizon typically lack strong theoretical guarantees such as orthogonality and quasi-oracle efficiency. In this paper, we revisit the problem of predicting individualized potential outcomes in sequential decision-making (i.e., estimating Q-functions in Markov decision processes with observational data) through a causal inference lens. In particular, we develop a comprehensive theoretical foundation for meta-learners in this setting with a focus on beneficial theoretical properties. As a result, we yield a novel meta-learner called DRQ-learner and establish that it is: (1) doubly robust (i.e., valid inference under the misspecification of one of the nuisances), (2) Neyman-orthogonal (i.e., insensitive to first-order estimation errors in the nuisance functions), and (3) achieves quasi-oracle efficiency (i.e., behaves asymptotically as if the ground-truth nuisance functions were known). Our DRQ-learner is applicable to settings with both discrete and continuous state spaces. Further, our DRQ-learner is flexible and can be used together with arbitrary machine learning models (e.g., neural networks). We validate our theoretical results through numerical experiments, thereby showing that our meta-learner outperforms state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。