arXiv:2411.07679cs.LGcs.GT2024-11NeurIPS被引 3

研究错误信念下博弈策略的收益风险权衡,给出理论边界。

Safe Exploitative Play with Untrusted Type Beliefs

  • 用贝叶斯博弈框架分析代理对其他智能体类型信念的依赖性
  • 发现错误信念导致的实际收益与最优收益之间存在可量化差距
  • 适用于需要鲁棒决策的多智能体系统,如自动驾驶协同

贝叶斯博弈与学习的结合历史悠久,其核心思想是:在多个行为未知的智能体组成的系统中,通过预设若干类型(每种类型代表一种可能行为),规划自身行动以应对最可能提高收益的类型。然而,这些类型信念通常基于过往行为学习而来,很可能不准确。本文关注一个智能体在面对其他组件的类型预测时,错误信念对其收益的影响。特别地,我们形式化定义了风险与机会之间的权衡,通过比较实际收益与最优收益之间的差距,揭示出信任或怀疑学习到的信念所导致的代价。主要结果为正常型和随机型贝叶斯博弈建立了帕累托前沿的上下界,并提供了数值验证。

原文摘要 · Abstract (English)

The combination of the Bayesian game and learning has a rich history, with the idea of controlling a single agent in a system composed of multiple agents with unknown behaviors given a set of types, each specifying a possible behavior for the other agents. The idea is to plan an agent's own actions with respect to those types which it believes are most likely to maximize the payoff. However, the type beliefs are often learned from past actions and likely to be incorrect. With this perspective in mind, we consider an agent in a game with type predictions of other components, and investigate the impact of incorrect beliefs to the agent's payoff. In particular, we formally define a tradeoff between risk and opportunity by comparing the payoff obtained against the optimal payoff, which is represented by a gap caused by trusting or distrusting the learned beliefs. Our main results characterize the tradeoff by establishing upper and lower bounds on the Pareto front for both normal-form and stochastic Bayesian games, with numerical results provided.

贝叶斯博弈多智能体鲁棒决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。