arXiv:2502.02020cs.LGstat.ME2025-02被引 1

在未知因果图下,用观测与实验数据结合优化干预选择

Causal Bandit Over Unknown Graphs: Upper Confidence Bounds With Backdoor Adjustment

  • 通过观测和实验数据识别后门调整集,融合信息构造置信上界
  • 理论证明算法累积后悔率更优,对干预臂数量依赖显著降低
  • 适用于有隐藏混杂因素的复杂系统,适合因果推断研究者

因果老虎机问题旨在通过序列实验,在由有向无环图(DAG)建模的因果系统中找出能最大化期望回报的干预措施。现有方法通常假设因果图已知或施加严格结构限制。本文研究因果图未知情形下的因果老虎机问题。首先考虑无隐藏混杂因子的高斯DAG模型,通过在老虎机过程中顺序收集的观测与实验数据,识别每个干预臂的候选后门调整集。这些集合可估计因果效应并构建融合两类数据信息的置信上界。基于此,提出新算法后门调整上界(BA-UCB),用于序列干预选择。我们建立了BA-UCB的有限时间累积后悔上界,显示其后悔率更优且对干预臂数量的依赖大幅减弱。进一步将方法与理论推广至存在隐藏混杂因子的情形,其中可观测变量由有向混合图建模。模拟实验表明,相较于现有方法,BA-UCB实现更低的累积后悔和更优的计算效率。

原文摘要 · Abstract (English)

The causal bandit problem seeks to identify, through sequential experimentation, an intervention that maximizes the expected reward in a causal system modeled by a directed acyclic graph (DAG). Existing methods typically assume that the causal graph is known or impose restrictive structural assumptions. In this paper, we study causal bandit problems when the causal graph is unknown. We first consider Gaussian DAG models without latent confounders. By combining observational and experimental data collected sequentially during the bandit process, we identify candidate backdoor adjustment sets for each intervention arm. These sets enable estimation of causal effects and construction of upper confidence bounds that integrate information from both data sources. Based on these estimates, we propose a new algorithm, termed backdoor-adjustment upper confidence bound (BA-UCB), for sequential intervention selection. We establish finite-time upper bounds on the cumulative regret of BA-UCB, showing improved rates and substantially relaxed dependency on the number of intervention arms compared to standard multi-armed bandit methods. We further extend the methodology and theoretical guarantees to settings with latent confounders, where the observed variables are modeled by an acyclic directed mixed graph. Simulation studies demonstrate that BA-UCB achieves substantially lower cumulative regret and favorable computational efficiency relative to existing approaches.

因果推断贝叶斯优化多臂老虎机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。