研究合作式多智能体中角色设定与实际协作模式的差距
Learned Coordination Conventions in Cooperative MARL: Measuring the Translation Gap Between Theory-Informed Roles and Learned Routing

- 用角色路由矩阵诊断智能体间协作结构
- 注意力机制使路由更集中且适应不同团队规模
- 适合关注多智能体协作可解释性的研究者
角色语义分配为异构智能体的协调方式提供了先验假设,但合作式多智能体强化学习系统通过去中心化、非平稳的学习过程形成协作惯例,无法保证其结果结构与先验一致。我们通过结合角色-路由矩阵、形态敏感性(Δ_max)以及梯度/遮蔽归因,考察三角色MiniGrid和SMACv2(Terran)环境中的这种理论预期与实际学习结构之间的翻译差距。结果表明,标签条件注意力相比扁平MLP基线产生更集中、更具角色特异性的路由,且在3v3至9v9规模下保持稳定,支持零样本跨团队规模迁移,并对友方槽填充不变。五次种子重评估显示,学习到的惯例与设计者指定先验存在部分对齐,但也揭示了小样本噪声可能制造出看似策略分歧的现象。本研究提出一个实证框架用于度量合作式多智能体中的协调结构,而非提出新的均衡概念或因果解释。
原文摘要 · Abstract (English)
Role-semantic assignments provide priors over how heterogeneous agents may coordinate, but cooperative MARL systems instead settle on conventions through decentralized, non-stationary learning, with no guarantee that the resulting structure matches those priors. We study this translation gap between theory-informed role expectations and learned coordination structure through a diagnostic combining a role-routing matrix, formation sensitivity ($Δ_{\max}$), and gradient/occlusion attribution across three-role MiniGrid and SMACv2 (Terran) environments. We show that label-conditioned attention produces substantially more concentrated and role-specific routing than flat MLP baselines, remains stable under 3v3--9v9 scaling, transfers zero-shot across team sizes, and is invariant to ally-slot padding. A 5-seed re-evaluation shows partial alignment between learned conventions and designer-specified priors while revealing where small-n noise can manufacture apparent strategic divergence. We present these results as an empirical framework for measuring coordination structure in cooperative MARL rather than as a new equilibrium concept or causal explanation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。