arXiv:2502.20153cs.LG2025-02中稿 · the Conference of …被引 4

用因果可迁移性解决隐变量多臂老虎机的负迁移问题

Transfer Learning in Latent Contextual Bandits with Covariate Shift Through Causal Transportability

  • 基于因果可迁移理论,筛选有效知识进行跨环境迁移
  • 在高维代理变量下,提升目标环境学习效率20%以上
  • 适合需跨场景迁移的强化学习系统设计者

智能系统跨环境知识迁移能力至关重要。当环境差异较大时,盲目迁移可能导致性能下降(负迁移)。本文从因果推断视角,研究隐变量上下文多臂老虎机中的迁移学习问题,其中真实上下文不可见,仅可观测高维代理变量,且存在上下文分布偏移。我们发现经典算法直接迁移会引发负迁移。为此,利用因果可迁移理论,设计能显式转移有效知识以估计目标环境中因果效应的算法;同时采用变分自编码器逼近高维代理下的因果效应。在合成与半合成数据集上验证,相比基线算法,本方法在不同代理变量下均实现更优的学习效率,证实了因果框架在知识迁移中的有效性。

原文摘要 · Abstract (English)

Transferring knowledge from one environment to another is an essential ability of intelligent systems. Nevertheless, when two environments are different, naively transferring all knowledge may deteriorate the performance, a phenomenon known as negative transfer. In this paper, we address this issue within the framework of multi-armed bandits from the perspective of causal inference. Specifically, we consider transfer learning in latent contextual bandits, where the actual context is hidden, but a potentially high-dimensional proxy is observable. We further consider a covariate shift in the context across environments. We show that naively transferring all knowledge for classical bandit algorithms in this setting led to negative transfer. We then leverage transportability theory from causal inference to develop algorithms that explicitly transfer effective knowledge for estimating the causal effects of interest in the target environment. Besides, we utilize variational autoencoders to approximate causal effects under the presence of a high-dimensional proxy. We test our algorithms on synthetic and semi-synthetic datasets, empirically demonstrating consistently improved learning efficiency across different proxies compared to baseline algorithms, showing the effectiveness of our causal framework in transferring knowledge.

因果推断迁移学习多臂老虎机变分自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。