在因果机制不确定时,用结构方程模型提升多臂老虎机的评估与学习效果。
Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms
- 基于结构方程模型处理因果关系不确定性,实现更稳健的策略评估。
- 相比传统方法,新方法在多种因果机制下评估更准确,且学习低方差最优策略。
- 适合需要高可靠性决策的场景,如医疗、金融等复杂系统优化。
因果图模型可编码来自领域专家背景知识或随机实验/观测数据发现的大量结构知识。然而,尽管我们可能了解因果关系的总体结构,却常不清楚确切的因果机制。本文提出一种能在条件概率分布不确定情况下有效推理的因果多臂老虎机评估与学习算法。进一步表明,条件独立性检验可用于选择建模变量。结果发现,结构方程模型(SEM)方法在评估精度上优于传统方法,尤其当可能的因果机制范围增大时。此外,该方法能学习低方差策略,并在模型充分正确指定时收敛至最优策略;而传统方法可能陷入局部极值或完全无法收敛。
原文摘要 · Abstract (English)
Causal graphical models can encode large amounts structural knowledge, both from the background knowledge of domain experts and the structural knowledge discovered from randomized experiments or observational data. However, though we may know the general structure of causal relationships, we often do not know the exact causal mechanisms. In this work, we propose a causal multi-armed bandit evaluation and learning algorithm that can reason effectively despite uncertainty over conditional probability distributions. Further, we show how conditional independence testing can be used to choose variables for modeling. We find that the structural equation model (SEM) approach gives more accurate evaluations compared to traditional approaches, particularly as the range of possible causal mechanisms grows. Further, the SEM approach learns low-variance policies, and it learns an optimal policy, assuming the model is sufficiently well-specified. Traditional approaches can converge to local extrema or fail to converge at all.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。