arXiv:2503.11452cs.LG2025-03

深度学习代理在避让游戏中自发演化出鹰鸽博弈行为。

Deep Learning Agents Trained For Avoidance Behave Like Hawks And Doves

  • 共享神经网络的双代理在对称网格中学习避让策略。
  • 一代理变激进抢占路径,另一代理学会规避以避免碰撞。
  • 行为模式类比动物界的鹰鸽博弈,适合研究智能体协作机制的人看。

我们提出由深度学习代理在简单避让游戏中表现的启发式最优策略。在对称网格世界中,两个代理必须交叉通过到达目标位置,同时避免彼此碰撞或偏离网格。两个代理共享同一个神经网络来决定策略。训练完成后,网络表现出类似鹰鸽博弈的行为:一个代理采用激进策略争夺路径,另一个则学会规避以避免冲突。

原文摘要 · Abstract (English)

We present heuristically optimal strategies expressed by deep learning agents playing a simple avoidance game. We analyse the learning and behaviour of two agents within a symmetrical grid world that must cross paths to reach a target destination without crashing into each other or straying off of the grid world in the wrong direction. The agent policy is determined by one neural network that is employed in both agents. Our findings indicate that the fully trained network exhibits behaviour similar to that of the game Hawks and Doves, in that one agent employs an aggressive strategy to reach the target while the other learns how to avoid the aggressive agent.

强化学习博弈行为多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。