提出受限信息共享的多智能体强化学习模型,解决隐私保护下的协作难题。
Learning with Limited Shared Information in Multi-agent Multi-armed Bandit
- 各智能体仅共享自愿信息,实现隐私保护下的协作
- 算法平均损失趋近常数,随智能体增多性能逼近最优
- 设计激励机制,确保参与协作的智能体都能获益
多智能体多臂赌博机(MAMAB)是一种经典的协同学习模型,近年来受到广泛关注。然而,现有研究未考虑某些智能体可能拒绝共享全部信息的情况,例如当部分数据涉及个人隐私时。本文提出一种新型的受限信息共享多智能体多臂赌博机(LSI-MAMAB)模型,其中每个智能体仅共享其愿意分享的信息,并提出平衡-探索与利用(Balanced-ETC)算法,以在信息受限条件下实现多智能体高效协作。理论分析表明,Balanced-ETC 算法渐近最优,当智能体数量充足时,每个智能体的平均累计遗憾趋于常数。此外,为鼓励智能体参与协作,设计了激励机制,确保每个智能体均能从系统中获益。最后,通过实验验证了理论结果的有效性。
原文摘要 · Abstract (English)
Multi-agent multi-armed bandit (MAMAB) is a classic collaborative learning model and has gained much attention in recent years. However, existing studies do not consider the case where an agent may refuse to share all her information with others, e.g., when some of the data contains personal privacy. In this paper, we propose a novel limited shared information multi-agent multi-armed bandit (LSI-MAMAB) model in which each agent only shares the information that she is willing to share, and propose the Balanced-ETC algorithm to help multiple agents collaborate efficiently with limited shared information. Our analysis shows that Balanced-ETC is asymptotically optimal and its average regret (on each agent) approaches a constant when there are sufficient agents involved. Moreover, to encourage agents to participate in this collaborative learning, an incentive mechanism is proposed to make sure each agent can benefit from the collaboration system. Finally, we present experimental results to validate our theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。