通过经验贝叶斯估计相关性,提升多任务强化学习的决策效率。
Empirical Bayesian Multi-Bandit Learning
- 构建分层贝叶斯框架,同时建模任务间相关性和差异性。
- 在真实和合成数据上,累积后悔值显著低于现有方法。
- 适合需跨任务共享信息的推荐系统与自适应控制场景。
上下文多臂老虎机中的多任务学习因其能通过共享结构和任务特异性增强决策而受到广泛关注。本文提出一种新的分层贝叶斯框架,用于学习多个带限制的老虎机实例。该框架通过分层贝叶斯模型捕捉不同老虎机实例间的异质性与相关性,实现有效信息共享的同时保留实例特异性。不同于以往忽略老虎机间协方差结构学习的方法,本文引入经验贝叶斯方法估计先验分布的协方差矩阵,提升了多老虎机学习的实用性与灵活性。基于此,我们设计了两种高效算法:ebmTS(经验贝叶斯多老虎机汤普森采样)和ebmUCB(经验贝叶斯多老虎机上置信界),均将估计的先验融入决策过程。我们为所提算法提供了频数论后悔上界,填补了多老虎机问题领域的研究空白。在合成与真实数据集上的大量实验表明,所提算法在复杂环境中表现优异,累积后悔显著低于现有技术,凸显其在探索与利用间平衡的有效性。
原文摘要 · Abstract (English)
Multi-task learning in contextual bandits has attracted significant research interest due to its potential to enhance decision-making across multiple related tasks by leveraging shared structures and task-specific heterogeneity. In this article, we propose a novel hierarchical Bayesian framework for learning in various bandit instances. This framework captures both the heterogeneity and the correlations among different bandit instances through a hierarchical Bayesian model, enabling effective information sharing while accommodating instance-specific variations. Unlike previous methods that overlook the learning of the covariance structure across bandits, we introduce an empirical Bayesian approach to estimate the covariance matrix of the prior distribution. This enhances both the practicality and flexibility of learning across multi-bandits. Building on this approach, we develop two efficient algorithms: ebmTS (Empirical Bayesian Multi-Bandit Thompson Sampling) and ebmUCB (Empirical Bayesian Multi-Bandit Upper Confidence Bound), both of which incorporate the estimated prior into the decision-making process. We provide the frequentist regret upper bounds for the proposed algorithms, thereby filling a research gap in the field of multi-bandit problems. Extensive experiments on both synthetic and real-world datasets demonstrate the superior performance of our algorithms, particularly in complex environments. Our methods achieve lower cumulative regret compared to existing techniques, highlighting their effectiveness in balancing exploration and exploitation across multi-bandits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。