Q-learning代理在重复博弈中会自发形成垄断定价,理论解释其行为机制。
Learning to Charge More: A Theoretical Study of Collusion by Q-Learning Agents
- 基于观测利润更新策略,无需计算均衡解
- 满足特定条件时持续收取高于竞争水平的价格
- 适用于研究智能体合谋的经济学与强化学习交叉领域
已有实验表明,Q-learninig代理可能学会收取超竞争性价格。本文首次在无限重复博弈中提供了理论解释:当博弈存在单阶段纳什均衡价格和可促成合谋的价格,且在实验结束时Q函数满足某些不等式时,企业会持续学习到收取超竞争性价格的行为。文章引入一类新的单记忆子博弈完美均衡(SPE),并给出学习行为由简单合谋、严惩触发或递增策略支持的条件。简单合谋仅在合谋启用价格即为单阶段纳什均衡时构成SPE,而严惩触发策略可以。
原文摘要 · Abstract (English)
There is growing experimental evidence that $Q$-learning agents may learn to charge supracompetitive prices. We provide the first theoretical explanation for this behavior in infinite repeated games. Firms update their pricing policies based solely on observed profits, without computing equilibrium strategies. We show that when the game admits both a one-stage Nash equilibrium price and a collusive-enabling price, and when the $Q$-function satisfies certain inequalities at the end of experimentation, firms learn to consistently charge supracompetitive prices. We introduce a new class of one-memory subgame perfect equilibria (SPEs) and provide conditions under which learned behavior is supported by naive collusion, grim trigger policies, or increasing strategies. Naive collusion does not constitute an SPE unless the collusive-enabling price is a one-stage Nash equilibrium, whereas grim trigger policies can.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。