提出可扩展的分布式智能体策略,优化最弱代理性能。
MEMOA: Massive Mixtures of Online Agents via Mean-Field Decentralized Nash Equilibria
- 基于均值场构建去中心化最优策略,仅需全局统计摘要
- 理论证明其渐近逼近中心化纳什最优,且在大规模下稳定收敛
- 在线加权机制提升整体预测精度,适合大规模分布式学习
在大规模AI时代,联邦学习虽重要但计算与通信成本难以随智能体数量增长而扩展。去中心化代理策略通过让每个智能体仅依赖自身状态和全局均值场实现自治。本文推导出闭式最优去中心化策略,以最小化最差代理的后悔(即最大在线损失)为优化目标。进一步证明该策略在大群体极限下渐近收敛至不可扩展的中心化纳什最优策略。通过在线加权机制优化服务器端客户端预测的混合结果,同时改进最弱代理与整体预测表现。数值实验验证了理论保证,并显示该策略普遍优于常见贪婪基线。
原文摘要 · Abstract (English)
In the modern age of large-scale AI, federated learning has become an increasingly important tool for training large populations of AI agents; however, its computational and communication costs can rapidly fail to scale with the number of agents. This is precisely where decentralized agentic strategies shine: each agent acts autonomously, using only its own state together with a minimal summary of the ensemble, namely the mean-field. We derive the unique optimal decentralized policy in closed form. Optimality is characterized through a worst-client/minimax criterion: minimizing the under-performer regret, namely the maximal online cost incurred by the weakest agent in the ensemble. We further prove that the resulting decentralized policy asymptotically converges, in the large-population limit, to the Nash-optimal centralized policy, whose direct computation is not scalable. We use an online weighting mechanism to optimize the server-computed mixture of client predictions, thereby improving the mean prediction in addition to the previously optimized weakest-client prediction. Numerical experiments verify our theoretical guarantees and demonstrate that our decentralized policy typically outperforms natural greedy decentralized baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。