解决多人协作抢资源时容量未知的难题,让系统自动学会合理分配。
Meet Me at the Arm: The Cooperative Multi-Armed Bandits Problem with Shareable Arms
- 设计协议驱动的分布式算法,通过试错发现每根资源杆的承载上限。
- 在容量有限且无法感知碰撞的条件下,实现对数级累积损失增长。
- 适合多智能体协同决策场景,如无线网络频谱共享或机器人任务分配。
我们研究在无感知条件下的去中心化多玩家多臂赌博机问题,每个玩家仅能获得自身奖励,无法获取碰撞信息。每根臂具有未知容量,若同时拉动该臂的玩家数量超过其容量,则所有参与者均获零回报。该设定推广了经典单位容量模型,在严重反馈限制下引入了协调与容量发现的新挑战。我们提出 A-CAPELLA(容量感知并行消除学习与分配算法),一种去中心化学习算法,通过协议驱动的协调机制,在此广义环境中实现对数级别遗憾。
原文摘要 · Abstract (English)
We study the decentralized multi-player multi-armed bandits (MMAB) problem under a no-sensing setting, where each player receives only their own reward and obtains no information about collisions. Each arm has an unknown capacity, and if the number of players pulling an arm exceeds its capacity, all players involved receive zero reward. This setting generalizes the classical unit-capacity model and introduces new challenges in coordination and capacity discovery under severe feedback limitations. We propose A-CAPELLA (Algorithm for Capacity-Aware Parallel Elimination for Learning and Allocation), a decentralized learning algorithm that achieves logarithmic regret in this generalized regime via protocol-driven coordination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。