研究多玩家信息不对称下的度量空间贝叶斯优化,提出新算法并证明其性能边界。
Multiplayer Information Asymmetric Bandits in Metric Spaces
- 采用固定与自适应离散化策略处理奖励和动作的信息不对称
- 在三种不对称场景下均获得与维度同阶的上界误差
- 适用于多智能体协同决策中存在信息差异的场景
近年来,信息不对称的利普希茨贝叶斯问题受到关注。本文研究了将利普希茨贝叶斯问题应用于多玩家信息不对称场景的问题,考虑奖励、动作或两者均存在信息不对称的情况。我们采用文献[1]提出的CAB算法,该算法使用固定离散化,在所有三种问题设置下均实现了与动作空间维度同阶的遗憾上界。同时,我们采纳文献[2]的zooming算法,该算法采用自适应离散化,将其应用于奖励信息不对称和动作信息不对称的场景,进一步提升效率。
原文摘要 · Abstract (English)
In recent years the information asymmetric Lipschitz bandits In this paper we studied the Lipschitz bandit problem applied to the multiplayer information asymmetric problem studied in \cite{chang2022online, chang2023optimal}. More specifically we consider information asymmetry in rewards, actions, or both. We adopt the CAB algorithm given in \cite{kleinberg2004nearly} which uses a fixed discretization to give regret bounds of the same order (in the dimension of the action) space in all 3 problem settings. We also adopt their zooming algorithm \cite{ kleinberg2008multi}which uses an adaptive discretization and apply it to information asymmetry in rewards and information asymmetry in actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。