利用跨环境因果不变性,提升智能体在不同场景下的学习效率。
On Transportability for Structural Causal Bandits
- 基于因果图结构融合多源数据,实现跨环境知识迁移。
- 算法达到次线性后悔率,且性能随先验数据信息量提升而增强。
- 适合有多个异构数据源的强化学习部署场景。
具备因果知识的智能体可通过利用潜在因果结构,识别无法最大化收益的动作,从而优化动作空间并避免不必要的探索。结构化因果老虎机框架通过图形化表征,使智能体能基于已有知识,在在线交互中估计某些动作的期望回报。然而,如何将来自不同条件(观测或实验)及异构环境的数据组合进行信息迁移,仍缺乏系统指导。本文研究了具有可迁移性的结构化因果老虎机,将源环境的先验知识融合以增强目标环境的学习效果。我们证明,利用跨环境不变性可一致地改善学习性能。所提出的老虎机算法实现了次线性后悔率,且其性能显式依赖于先验数据的信息量,可能优于仅依赖在线学习的标准方法。
原文摘要 · Abstract (English)
Intelligent agents equipped with causal knowledge can optimize their action spaces to avoid unnecessary exploration. The structural causal bandit framework provides a graphical characterization for identifying actions that are unable to maximize rewards by leveraging prior knowledge of the underlying causal structure. While such knowledge enables an agent to estimate the expected rewards of certain actions based on others in online interactions, there has been little guidance on how to transfer information inferred from arbitrary combinations of datasets collected under different conditions -- observational or experimental -- and from heterogeneous environments. In this paper, we investigate the structural causal bandit with transportability, where priors from the source environments are fused to enhance learning in the deployment setting. We demonstrate that it is possible to exploit invariances across environments to consistently improve learning. The resulting bandit algorithm achieves a sub-linear regret bound with an explicit dependence on informativeness of prior data, and it may outperform standard bandit approaches that rely solely on online learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。