用因果方法提升离线决策学习的样本效率与可靠性
Learning Optimal and Sample-Efficient Decision Policies with Guarantees
- 基于工具变量和双重机器学习,解决隐藏混淆因子下的离线策略学习
- 算法在合成与半合成数据上显著优于现有方法,具备收敛与最优性保证
- 适合高风险场景如医疗、金融中的安全决策系统开发
强化学习与深度学习革新了决策范式,但在高风险应用中仍面临挑战,尤其在需要可证明保证的场景下。传统RL依赖大量在线交互,而离线学习则受隐藏混淆因子影响,导致虚假相关性与次优行为。本文首先利用工具变量(IV)识别因果效应,将问题建模为条件矩约束(CMR),借鉴双重/去偏机器学习思想,提出一种具有收敛性和最优性保证的样本高效算法。其次,在模仿学习中放宽对混淆因子的假设,将该估计器扩展为具收敛率保证的有效模仿策略学习算法。最后,针对线性时序逻辑(LTL)表达的高层目标,设计出可证明最优且样本高效的求解算法。通过在强化学习基准及合成与半合成数据集上的评估,验证了方法在真实决策场景中的有效性。
原文摘要 · Abstract (English)
The paradigm of decision-making has been revolutionised by reinforcement learning and deep learning. Although this has led to significant progress in domains such as robotics, healthcare, and finance, the use of RL in practice is challenging, particularly when learning decision policies in high-stakes applications that may require guarantees. Traditional RL algorithms rely on a large number of online interactions with the environment, which is problematic in scenarios where online interactions are costly, dangerous, or infeasible. However, learning from offline datasets is hindered by the presence of hidden confounders. Such confounders can cause spurious correlations in the dataset and can mislead the agent into taking suboptimal or adversarial actions. Firstly, we address the problem of learning from offline datasets in the presence of hidden confounders. We work with instrumental variables (IVs) to identify the causal effect, which is an instance of a conditional moment restrictions (CMR) problem. Inspired by double/debiased machine learning, we derive a sample-efficient algorithm for solving CMR problems with convergence and optimality guarantees, which outperforms state-of-the-art algorithms. Secondly, we relax the conditions on the hidden confounders in the setting of (offline) imitation learning, and adapt our CMR estimator to derive an algorithm that can learn effective imitator policies with convergence rate guarantees. Finally, we consider the problem of learning high-level objectives expressed in linear temporal logic (LTL) and develop a provably optimal learning algorithm that improves sample efficiency over existing methods. Through evaluation on reinforcement learning benchmarks and synthetic and semi-synthetic datasets, we demonstrate the usefulness of the methods developed in this thesis in real-world decision making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。