arXiv:2411.13690cs.LG2024-11

多智能体协作在线性伯努利模型中高效识别最优选项

Multi-Agent Best Arm Identification in Stochastic Linear Bandits

  • 通过中心服务器共享信息,多智能体并行探索并逐步淘汰差选项
  • 错误率随预算增长呈指数下降,理论与实验证明优于现有方法
  • 适合分布式决策场景,如智能推荐、资源调度等协同任务

我们研究在固定预算下,通过星型网络连接的多个智能体在随机线性贝叶斯问题中协同识别最优臂。智能体并行与线性贝叶斯实例交互,通过中心服务器共享知识,目标是最大限度降低最优臂估计的错误概率。为此,提出两种算法:适用于星型拓扑的MaLinBAI-Star和适用于任意拓扑的MaLinBAI-Gen。两者结合G-最优设计与逐轮剔除策略,在每轮通信中交换信息。理论上和实验上均证明,错误概率随时间预算呈指数衰减。在合成数据和真实世界数据上的实验表明,所提算法在性能上显著优于现有最先进多智能体算法。

原文摘要 · Abstract (English)

We study the problem of collaborative best-arm identification in stochastic linear bandits under a fixed-budget scenario. In our learning model, we first consider multiple agents connected through a star network, interacting with a linear bandit instance in parallel. We then extend our analysis to arbitrary network topologies. The objective of the agents is to collaboratively identify the best arm of the given bandit instance with the help of a central server while minimizing the probability of error in best arm estimation. To this end, we propose two algorithms, MaLinBAI-Star and MaLinBAI-Gen for star networks and networks with arbitrary structure, respectively. Both algorithms utilize the technique of G-optimal design along with the successive elimination based strategy where agents share their knowledge through a central server at each communication round. We demonstrate, both theoretically and empirically, that our algorithms achieve exponentially decaying probability of error in the allocated time budget. Furthermore, experimental results on both synthetic and real-world data validate the effectiveness of our algorithms over the state-of-the art existing multi-agent algorithms.

多智能体在线学习最佳臂识别线性带宽

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。