arXiv:2605.23930cs.AIcs.LG2026-05

用时间量化机制研究双智能体协作,发现同步冲关最有效

Quantum Frog: Emergent Cooperation and Difficulty Scaling in a Quantized-Time Cooperative Game

论文配图:Quantum Frog: Emergent Cooperation and Difficulty Scaling in a Quantized-Time Cooperative Game
图 1 · 摘自论文原文
  • 环境仅在玩家行动时推进,强制时间敏感协作
  • 合作训练使成功率提升32%~34%,平均步数从90减至6
  • 共享目标即可实现同步冲刺,无需复杂协调

我们提出 extit{Quantum Frog},一款基于新颖时间量化机制的双人协作游戏,环境仅在玩家行动时推进。受经典游戏Frogger启发,两名青蛙需共同穿越8×8交通网格。通过强化学习分析四个设计问题:(1)难度如何随车流密度变化,(2)单智能体最优策略为何,(3)独立与协作双智能体间的协作差距多大,(4)激励合作后涌现何种联合策略。训练经历五阶段:表格式Q-learning、深度Q网络(DQN)、独立DQN(IDQN)、带中心化评价器的多智能体近端策略优化(MAPPO),评估车流密度为1至6时的表现。关键发现:(i)时间量化机制使“冲关策略”(每步直线上行)成为普遍最优,因能最小化暴露于交通的时间;(ii)加入未协同的第二名玩家比单个专家玩家面对六倍车流更难;(iii)合作训练相较独立训练提升32%~34%的联合成功率,并将平均回合长度从~90步缩短至~6步;(iv)涌现的协作策略为同步冲关,而非复杂位置协调,表明在时间敏感任务中,共享激励足以对齐智能体行为。这些发现为 extit{Quantum Frog}的商业设计提供实证指导,并深化了对环境机制影响多智能体学习动态的理解。

原文摘要 · Abstract (English)

We introduce \emph{Quantum Frog}, a two-player cooperative game built on a novel \emph{quantized-time} mechanic in which the environment advances only when a player acts. Inspired by the classic arcade game Frogger, Quantum Frog requires two frogs to cross an 8$\times$8 grid of traffic and reach the far side together. We use reinforcement learning (RL) as an analytical lens to answer four design questions: (1) how does game difficulty scale with traffic density, (2) what is the optimal single-agent policy and why, (3) how large is the cooperation gap between independent and cooperative two-agent play, and (4) what joint strategy emerges when agents are incentivised to cooperate? We train agents through five escalating stages, Tabular Q-Learning, Deep Q-Network (\DQN), Independent \DQN~(\IDQN), and Multi-Agent Proximal Policy Optimisation (\MAPPO\ with a centralised critic), evaluating each against traffic densities of one to six cars. Our key findings are: (i) the quantized-time mechanic makes a \emph{rush strategy} (moving directly upward at every step) universally optimal, as time exposure to traffic is minimised; (ii) adding an uncoordinated second player is harder than sextupling the traffic for a single expert player; (iii) cooperative training recovers +32--34 percentage points of joint success rate relative to independent agents and reduces episode length from $\sim$90 to $\sim$6 steps; and (iv) the emergent cooperative strategy is synchronised rushing, not complex positional coordination, illustrating that shared incentives alone suffice to align agents in time-critical cooperative tasks. These findings provide concrete, empirically grounded guidance for the commercial design of Quantum Frog and offer broader insights into the role of environment mechanics in shaping multi-agent learning dynamics.

多智能体协作强化学习游戏机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。