arXiv:2607.15577cs.LGcs.AI2026-07

利用因果结构共享信息,提升无法直接干预变量的决策效率。

Information-Directed Sampling for Causal Bandits

论文配图:Information-Directed Sampling for Causal Bandits
图 1 · 摘自论文原文
  • 基于贝叶斯框架建模因果机制,通过观测更新跨干预的奖励估计。
  • 提出因果版信息导向采样算法,在合成任务中显著优于基线方法。
  • 适用于有不可控变量的决策场景,如医疗、推荐系统等复杂系统优化。

因果老虎机利用变量间的结构性关系,在不同干预之间共享信息,加速高回报决策的识别。但在许多应用中,某些变量无法直接操控,尽管它们影响收益并提供关于潜在因果系统的有用信息。本文研究带有非可控变量的上下文因果老虎机,其中上下文变量在动作选择前被观测,额外变量在每次干预后被观测。假设已知无潜混杂的因果图,采用贝叶斯形式化,将观测分布的条件概率表作为未知参数。该表示允许某一干预下的观测结果通过共享的因果机制更新其他干预的收益估计。我们开发了因果版的汤普森采样和信息导向采样(IDS)算法。对于汤普森采样,建立了依赖熵的亚线性贝叶斯后悔界;对于IDS,推导出依赖熵的后悔界,显式量化了蒙特卡洛近似引入的额外误差;当这些量可精确获得时,该界恢复标准的亚线性IDS速率。进一步提供了算法中使用的蒙特卡洛估计的高概率置信界。在多个合成因果老虎机任务上的实验表明,所提方法通过更有效地利用干预间的共享信息,显著优于因果与非因果基线。

原文摘要 · Abstract (English)

Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly manipulated, even though they influence the reward and provide useful information about the underlying causal system. We study contextual causal bandits with non-manipulable variables, where context variables are observed before action selection and additional variables are observed after each intervention. Assuming a known causal graph without latent confounding, we adopt a Bayesian formulation in which the conditional probability tables of the observational distribution constitute the unknown parameter. This representation allows observations collected under one intervention to update reward estimates for other interventions through their shared causal mechanisms. We develop causal variants of Thompson Sampling and Information-Directed Sampling (IDS) for this setting. For Thompson Sampling, we establish an entropy-dependent sublinear Bayesian regret bound. For IDS, we derive an entropy-dependent regret bound that explicitly quantifies the additional error introduced by Monte Carlo approximation of the expected regret and information gain; when these quantities are available exactly, the bound recovers the standard sublinear IDS rate. We further provide high-probability confidence bounds for the Monte Carlo estimates used by the algorithm. Experiments on several synthetic causal bandit tasks show that the proposed methods outperform causal and non-causal baselines by more effectively exploiting information shared across interventions.

因果推断强化学习贝叶斯优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。