arXiv:2509.20147cs.GTcs.LG2025-09中稿 · publication at IEE…被引 1

多游戏环境下,玩家通过低通信量策略实现高效协作与均衡分配。

Choose Your Battles: Distributed Learning Over Multiple Tug of War Games

  • 用随机近似更新动作,1位通信决定游戏切换
  • 算法收敛至满足服务质量目标的均衡点
  • 适用于传感器网络、任务分配等分布式场景

考虑N名玩家和K个同时进行的博弈,每个博弈建模为拉力赛(Tug-of-War, ToW)游戏,其中一名玩家增加行动会降低其他所有玩家的收益。每位玩家在任意时刻仅参与一个游戏。每一步中,玩家选择参与的游戏及采取的动作,其收益取决于同游戏内所有玩家的行为。这一由K个游戏构成的系统称为“元拉力赛”(Meta-ToW)游戏,可模拟功率控制、分布式任务分配及传感器网络激活等场景。本文提出元拉力赛和平算法(Meta Tug-of-Peace),一种分布式算法:动作更新采用简单随机逼近方法,游戏切换决策通过玩家间稀疏的1比特通信完成。证明了该算法在Meta-ToW游戏中能收敛至满足目标服务质量收益向量的均衡点。并通过模拟验证了其在上述场景中的有效性。

原文摘要 · Abstract (English)

Consider $N$ players and $K$ games taking place simultaneously. Each of these games is modeled as a Tug-of-War (ToW) game where increasing the action of one player decreases the reward for all other players. Each player participates in only one game at any given time. At each time step, a player decides the game in which they wish to participate in and the action they take in that game. Their reward depends on the actions of all players that are in the same game. This system of $K$ games is termed a 'Meta Tug-of-War' (Meta-ToW) game. These games can model scenarios such as power control, distributed task allocation, and activation in sensor networks. We propose the Meta Tug-of-Peace algorithm, a distributed algorithm where the action updates are done using a simple stochastic approximation algorithm, and the decision to switch games is made using an infrequent 1-bit communication between the players. We prove that in Meta-ToW games, our algorithm converges to an equilibrium that satisfies a target Quality of Service reward vector for the players. We then demonstrate the efficacy of our algorithm through simulations for the scenarios mentioned above.

分布式学习博弈论通信优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。