提出机制让自私代理在多臂赌博机中持续协作,靠信息共享而非金钱奖励。
Collaborating in Multi-Armed Bandits with Strategic Agents
- 设计CAOS机制,让自私代理因长期参与而自愿共享信息。
- 在战略行为下仍保持接近完全合作系统的性能,后悔值接近最优。
- 适合研究激励机制与多智能体协作的学者和工程师。
我们研究多智能体贝叶斯多臂赌博机中的协同学习问题,其中多个策略性代理共同解决同一赌博机实例。虽然信息共享可加速学习,但策略性代理可能倾向于搭便车、规避探索。本文考虑长期参与的持久代理(即在多个时间周期中持续参与),不同于以往多数研究中假设的短时代理(每个代理仅做一次决策)。在无货币转移、仅通过信息共享提供激励的设定下,我们提出 exttt{CAOS}机制,可在纳什均衡下维持协作,并实现强后悔界。结果表明,仅依靠信息共享即可维持协同探索,性能接近完全合作系统,即使存在策略性行为。
原文摘要 · Abstract (English)
We study collaborative learning in multi-agent Bayesian bandit problems, where strategic agents collectively solve the same bandit instance. While multiple agents can accelerate learning by sharing information, strategic agents might prefer to free-ride and avoid exploration. We consider a setting with persistent agents that participate in multiple time periods. This is in contrast to most previous works on incentives in multi-agent MAB, which assume short-lived agents, namely each agent has a single decision to make and optimizes their expected reward in that single decision. As in the multi-agent MAB model with incentives, our model does not have monetary transfers, and the only incentives are through information sharing. We propose \texttt{CAOS}, a mechanism that sustains collaboration as a Nash equilibrium while achieving strong regret guarantees. Our results demonstrate that collaborative exploration can be sustained purely through information sharing, achieving performance close to that of fully cooperative systems despite strategic behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。