arXiv:2512.12693cs.LGcs.AI2025-12

通过共享结构实现多任务强化学习中的高效探索与利用

Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits

  • 基于贝叶斯框架,利用任务间潜在依赖关系整合全局观测
  • 在部分上下文观测下,仍能有效降低结构与用户特定不确定性
  • 适用于存在复杂异质性或模型不匹配的多任务场景

我们提出一种新颖的贝叶斯框架,用于情境化多任务多臂赌博机设置中的高效探索,其中仅可部分观测上下文,且奖励分布间的依赖关系由潜在上下文变量诱导。为利用这些结构依赖,该方法整合所有任务的观测并学习全局联合分布,同时仍支持对新任务的个性化推断。我们识别出两类关键认知不确定性:跨任务和跨臂的潜在奖励依赖结构不确定性,以及因上下文不完整和交互历史有限带来的用户特定不确定性。为付诸实践,我们使用基于粒子的对数密度高斯过程表示任务与奖励的联合分布,实现无需预设假设即可灵活、数据驱动地发现臂间与任务间依赖关系。实验表明,本方法在模型误设或复杂潜在异质性场景中显著优于层次化模型赌博机等基线方法。

原文摘要 · Abstract (English)

We propose a novel Bayesian framework for efficient exploration in contextual multi-task multi-armed bandit settings, where the context is only observed partially and dependencies between reward distributions are induced by latent context variables. In order to exploit these structural dependencies, our approach integrates observations across all tasks and learns a global joint distribution, while still allowing personalised inference for new tasks. In this regard, we identify two key sources of epistemic uncertainty, namely structural uncertainty in the latent reward dependencies across arms and tasks, and user-specific uncertainty due to incomplete context and limited interaction history. To put our method into practice, we represent the joint distribution over tasks and rewards using a particle-based approximation of a log-density Gaussian process. This representation enables flexible, data-driven discovery of both inter-arm and inter-task dependencies without prior assumptions on the latent variables. Empirically, we demonstrate that our method outperforms baselines such as hierarchical model bandits, especially in settings with model misspecification or complex latent heterogeneity.

多任务学习贝叶斯优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。