将有限理性引入平均场博弈,更真实地模拟大规模群体行为。
Bounded Rationality Equilibrium Learning in Mean Field Games
- 用量化响应均衡建模个体对目标的噪声估计
- 引入滚动时域机制限制决策规划范围
- 适合研究非完全理性的群体智能系统
平均场博弈(MFG)可有效建模大规模群体行为。现有学习方法多聚焦于纳什均衡(NE),假设个体完全理性,但在现实中常不成立。为此,本文引入有限理性概念,基于量化响应均衡(QRE)构建两类新型MFG QRE,以刻画个体仅能噪声式估计真实目标的情形。同时,通过限制代理的规划时域,提出新型滚动时域(RH)MFG,结合QRE与已有方法,综合建模多种有限理性特征。论文形式化定义了MFG QRE与RH MFG,并与熵正则化纳什均衡等概念对比。进一步设计广义不动点迭代与虚构博弈算法用于学习QRE与RH均衡。经理论分析与多组实验验证,展示了算法能力,并阐明不同均衡概念的实际差异。
原文摘要 · Abstract (English)
Mean field games (MFGs) tractably model behavior in large agent populations. The literature on learning MFG equilibria typically focuses on finding Nash equilibria (NE), which assume perfectly rational agents and are hence implausible in many realistic situations. To overcome these limitations, we incorporate bounded rationality into MFGs by leveraging the well-known concept of quantal response equilibria (QRE). Two novel types of MFG QRE enable the modeling of large agent populations where individuals only noisily estimate the true objective. We also introduce a second source of bounded rationality to MFGs by restricting agents' planning horizon. The resulting novel receding horizon (RH) MFGs are combined with QRE and existing approaches to model different aspects of bounded rationality in MFGs. We formally define MFG QRE and RH MFGs and compare them to existing equilibrium concepts such as entropy-regularized NE. Subsequently, we design generalized fixed point iteration and fictitious play algorithms to learn QRE and RH equilibria. After a theoretical analysis, we give different examples to evaluate the capabilities of our learning algorithms and outline practical differences between the equilibrium concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。