arXiv:2607.12861cs.ROcs.AI2026-07中稿 · IROS 2026

用新工具揭示机器人集群如何从简单奖励中自发形成复杂行为

Unveiling Complex Collective Behaviors from Simple Rewards

论文配图:Unveiling Complex Collective Behaviors from Simple Rewards
图 1 · 摘自论文原文
  • 提出双阶段解释框架与响应地图,分析机器人决策空间模式
  • 发现机器人自动利用环境几何结构作为导航目标,无需显式指令
  • 适用于研究集群智能、强化学习可解释性,适合机器人与AI交叉领域

多智能体强化学习(MARL)在机器人集群中潜力巨大,但神经网络策略的黑箱特性阻碍了战略分析,限制了实际应用。更令人意外的是,复杂集群行为可从简单奖励中自发涌现,而无需显式的聚合激励。揭示这一过程背后的机制至关重要,但简单奖励与集体行为之间的脱节加剧了可解释性挑战。本文旨在揭示此过程中的隐藏机制,提出一种两阶段的EEC( LinkIII)解释框架,并引入一种新型分析工具——代理响应图(ARM),用于揭示代理在空间中的决策模式,识别聚集与回避区域。实验表明,机器人隐式学习了环境的几何场,并将其作为协调运动的目标。在协作任务中,ARM识别出未被占据的目标内部区域为导航目的地;当中心被占据时,该目标自动转向边界,体现自主探索未占用区域的能力。在竞争任务中,ARM意外发现捕食者沃罗诺伊图边界是猎物汇聚的终点。两个任务共同验证了ARM在揭示机器人集群MARL策略下隐藏几何结构方面的有效性。

原文摘要 · Abstract (English)

Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic analysis, limiting multi-robot applications. Furthermore, complex swarm behaviors can surprisingly emerge from simple rewards without explicit aggregation incentives. Unveiling the mechanisms behind this emergence is critical, but the disconnection between simple rewards and collective behaviors exacerbates interpretability challenges. This paper aims to reveal the hidden mechanisms in this process. We propose a two-stage EEC (\LinkIII) explanatory framework. This includes a novel analytical tool called the Agent Response Map (ARM), which reveals agents' decision-making patterns across space and identifies regions of aggregation and avoidance. ARM reveals that the robots implicitly learn the geometric fields of the environment and utilize these structures as desired targets for coordinated movement. We validate this finding across two distinct tasks: a cooperative multi-robot shape assembly and a competitive predator-prey pursuit-evasion. 1) In the cooperative task, ARM identifies the unoccupied target interior as the desired destination for robot navigation. As the center becomes occupied, this target region automatically shifts toward the boundary, demonstrating the robots' capacity to autonomously explore unoccupied areas. 2) In the competitive task, ARM surprisingly identifies the boundary of the predators' Voronoi diagram as the convergence destination for prey agents. Together, these two tasks demonstrate the capability of ARM to discover the hidden geometric structures underlying MARL policies in robot swarms.

多智能体强化学习可解释性集群智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。