用多智能体强化学习研究交易算法如何自发形成合谋或竞争。
Multi-Agent Reinforcement Learning for Market Making: Competition without Collusion
- 构建分层框架,让不同目标的交易代理相互博弈。
- 竞争型代理使买卖价差缩小,提升市场执行效率。
- 自适应代理能共存且对他人收益影响较小,更可持续。
算法合谋已成为人工智能领域核心问题:部署于市场的多个智能体交互是否会导致合谋?更广泛地,理解涌现行为(如卡特尔或高级机器人主导市场)对整体市场的影响至关重要。本文提出一种分层多智能体强化学习框架,用于研究做市中的算法合谋。框架包含一个自利做市商(代理A),其在对手塑造的不确定环境中训练;以及三个底层竞争者:自利代理B1(最大化自身盈亏)、竞争代理B2(最小化对手盈亏)和混合代理B*,可动态切换前两者的行为。为分析代理间相互作用及市场结果,我们设计了交互级度量指标,量化行为不对称性与系统级动态,提供潜在的涌现模式信号。实验表明,在零和设置中,代理B2击败代理B1,激进捕获订单流并压缩平均价差,从而提升市场执行效率。而代理B*在与其他盈利驱动代理共存时表现出自利倾向,通过自适应报价获取主导市场份额,但对代理A和B1的收益负面影响较B2温和。这些发现表明,自适应激励控制有助于异构智能体环境中的可持续战略共存,并为算法交易系统的行为设计提供了结构化评估视角。
原文摘要 · Abstract (English)
Algorithmic collusion has emerged as a central question in AI: Will the interaction between different AI agents deployed in markets lead to collusion? More generally, understanding how emergent behavior, be it a cartel or market dominance from more advanced bots, affects the market overall is an important research question. We propose a hierarchical multi-agent reinforcement learning framework to study algorithmic collusion in market making. The framework includes a self-interested market maker (Agent~A), which is trained in an uncertain environment shaped by an adversary, and three bottom-layer competitors: the self-interested Agent~B1 (whose objective is to maximize its own PnL), the competitive Agent~B2 (whose objective is to minimize the PnL of its opponent), and the hybrid Agent~B$^\star$, which can modulate between the behavior of the other two. To analyze how these agents shape the behavior of each other and affect market outcomes, we propose interaction-level metrics that quantify behavioral asymmetry and system-level dynamics, while providing signals potentially indicative of emergent interaction patterns. Experimental results show that Agent~B2 secures dominant performance in a zero-sum setting against B1, aggressively capturing order flow while tightening average spreads, thus improving market execution efficiency. In contrast, Agent~B$^\star$ exhibits a self-interested inclination when co-existing with other profit-seeking agents, securing dominant market share through adaptive quoting, yet exerting a milder adverse impact on the rewards of Agents~A and B1 compared to B2. These findings suggest that adaptive incentive control supports more sustainable strategic co-existence in heterogeneous agent environments and offers a structured lens for evaluating behavioral design in algorithmic trading systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。