通过相对表示发现多智能体强协作选项,提升协同效率。
Inter-Agent Relative Representations for Multi-Agent Option Discovery
- 构建团队状态对齐的抽象表征,捕捉智能体间同步模式。
- 在多个仿真场景中,新方法生成的选项使下游协作能力更强。
- 适合关注多智能体协作与选项发现的研究者。
时序扩展动作在单智能体设置中能增强探索与规划能力。在多智能体场景中,联合状态空间随智能体数量呈指数增长,使得协调行为更加重要。然而,这一特性也使多智能体选项设计尤为困难。现有方法常因产生松散耦合或完全独立的行为而牺牲协调性。为此,本文提出一种新型多智能体选项发现方法:构建联合状态抽象,在压缩状态空间的同时保留发现强协调行为所需信息。该方法基于一个归纳偏置——智能体状态间的同步可自然构成无显式目标下的协调基础。首先近似团队最大对齐的虚构状态(即“Fermat”状态),并以此定义衡量每个状态维度上团队层面偏离程度的‘spreadness’指标。基于此表示,进一步使用神经图拉普拉斯估计器推导出捕捉智能体间状态同步模式的选项。我们在两个模拟多智能体领域中的多个场景下评估了所得选项,结果表明其相比其他选项发现方法具有更强的下游协调能力。
原文摘要 · Abstract (English)
Temporally extended actions improve the ability to explore and plan in single-agent settings. In multi-agent settings, the exponential growth of the joint state space with the number of agents makes coordinated behaviours even more valuable. Yet, this same exponential growth renders the design of multi-agent options particularly challenging. Existing multi-agent option discovery methods often sacrifice coordination by producing loosely coupled or fully independent behaviours. Toward addressing these limitations, we describe a novel approach for multi-agent option discovery. Specifically, we propose a joint-state abstraction that compresses the state space while preserving the information necessary to discover strongly coordinated behaviours. Our approach builds on the inductive bias that synchronisation over agent states provides a natural foundation for coordination in the absence of explicit objectives. We first approximate a fictitious state of maximal alignment with the team, the \textit{Fermat} state, and use it to define a measure of \textit{spreadness}, capturing team-level misalignment on each individual state dimension. Building on this representation, we then employ a neural graph Laplacian estimator to derive options that capture state synchronisation patterns between agents. We evaluate the resulting options across multiple scenarios in two simulated multi-agent domains, showing that they yield stronger downstream coordination capabilities compared to alternative option discovery methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。