针对重尾奖励与信息不对称的多智能体博弈,提出鲁棒分布式算法。
Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

- 设计三种信息不对称下的鲁棒分布式算法
- 理论证明近似达到集中式重尾最优率
- 适合研究分布式决策与非高斯奖励场景
多臂赌博机是序列决策的核心框架,传统研究多基于次高斯奖励假设。然而真实应用中常面临重尾奖励分布和去中心化、信息不对称的交互。本文研究在三种信息不对称情形下的多智能体多臂赌博机:未观测动作但共享奖励、可观测动作且奖励独立、未观测动作且奖励独立。针对每种情形设计鲁棒分布式算法,并推导出近乎匹配集中式重尾最优率的泛化误差界。在帕累托分布奖励环境上的实验验证了理论结果,展示了同步、协调与探索之间的权衡关系。
原文摘要 · Abstract (English)
The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. However, real-world applications often involve heavy-tailed reward distributions and decentralized, information-asymmetric interactions. We study multi-agent multi-armed bandits with heavy-tailed rewards under three information-asymmetry regimes: unobserved actions with common rewards, observed actions with independent rewards, and unobserved actions with independent rewards. We develop robust decentralized algorithms for each setting and derive regret guarantees that nearly match centralized heavy-tailed rates. Experiments on a Pareto-distributed reward environment validate our theoretical findings and illustrate the trade-offs between synchronization, coordination, and exploration across the three regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。