arXiv:2606.09884cs.MAcs.AI2026-06

揭示深度多智能体强化学习在异步定价中的两种失效模式及修复方案。

Failure Modes of Deep Multi-Agent RL in Asynchronous Pricing: Reproducible Triggers, Trace Diagnostics, and a Partial Fix

  • 构建连续时间多智能体基准,模拟竞争性定价环境。
  • 同步训练导致代理合谋(合谋指数Δ=0.69±0.11),异步+延迟可降低48%。
  • 提供轨迹级诊断工具,揭示信号崩溃与冲击后无法恢复现象。

我们研究了深度多智能体强化学习在连续时间定价市场中的两种可复现的失效模式:(i) 竞争性DDPG智能体间隐性形成卡特尔;(ii) 高事件率下的演员-评论家不稳定。在单一连续时间多智能体强化学习基准(泊松时钟定价更新、观测延迟δ、内部最优逻辑斯蒂需求)中,我们验证了同步DDPG智能体在合谋指数Δ=0.69±0.11时稳定触发第一类失效,并量化了一种部分微观结构修复方案:仅异步即可使合谋减少48%,加入延迟更将合谋降至最低Δ=0.28。该修复有明确代价:非完全(Δ仍高于伯特兰水平)、δ上非单调,且无法应对第二类失效——当λ=5时评论家发散,破坏(λ=5, δ=1)区域相图。我们辅以轨迹级追踪诊断,揭示每回合内信号崩溃及冲击后无法恢复现象。

原文摘要 · Abstract (English)

We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation between competing DDPG agents, and (ii) actor--critic instability at high event rates. We instantiate both inside a single CT-MARL benchmark (Poisson-clocked price updates, observation latency $δ$, interior-optimum logit demand), show that synchronous DDPG agents reliably trigger Failure Mode 1 with collusion index $Δ= 0.69 \pm 0.11$, and quantify a partial microstructure fix: asynchrony alone cuts collusion by 48\% and adding latency drives it to a minimum of $Δ= 0.28$. The fix has clearly documented costs: it is partial ($Δ$ remains supra-Bertrand), it is non-monotone in $δ$, and it does not survive Failure Mode 2, which emerges as DDPG critic divergence at $λ= 5$ and corrupts the phase-diagram cell at $(λ{=}5, δ{=}1)$. We accompany the scalar collusion index with trajectory-level trace diagnostics that expose the within-episode signalling collapse and the post-shock non-recovery.

多智能体强化学习定价模型失效分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。