通过分析策略分布,量化多智能体中每个智能体的行为影响。
Understanding Action Effects through Instrumental Empowerment in Multi-Agent Reinforcement Learning
- 基于信息论的沙普利值,衡量智能体对同伴决策确定性与策略一致性的影响力。
- 在合作与竞争任务中,识别出促进团队成功的关键行为模式。
- 适合关注多智能体系统可解释性与协作机制的研究者。
为可靠部署多智能体强化学习(MARL)系统,理解个体智能体行为至关重要。以往工作通常依赖显式奖励信号评估团队整体表现,但在缺乏价值反馈时,难以推断智能体贡献。本文提出一种新方法:仅通过分析策略分布来获取有意义的行为洞察。受智能体趋向追求共同工具性价值现象启发,引入意图协作值(ICVs),基于信息论沙普利值量化每个智能体对其同伴工具性赋能的因果影响。具体而言,ICVs通过评估同伴决策的不确定性和偏好一致性,衡量智能体行为对队友策略的影响。在合作与竞争型MARL任务中,该方法揭示了哪些行为有助于团队成功——无论是促进确定性决策还是保留未来选择灵活性,并进一步反映智能体采用相似或多样策略的程度。本方法为理解协作动态提供了新视角,提升了MARL系统的可解释性。
原文摘要 · Abstract (English)
To reliably deploy Multi-Agent Reinforcement Learning (MARL) systems, it is crucial to understand individual agent behaviors. While prior work typically evaluates overall team performance based on explicit reward signals, it is unclear how to infer agent contributions in the absence of any value feedback. In this work, we investigate whether meaningful insights into agent behaviors can be extracted solely by analyzing the policy distribution. Inspired by the phenomenon that intelligent agents tend to pursue convergent instrumental values, we introduce Intended Cooperation Values (ICVs), a method based on information-theoretic Shapley values for quantifying each agent's causal influence on their co-players' instrumental empowerment. Specifically, ICVs measure an agent's action effect on its teammates' policies by assessing their decision (un)certainty and preference alignment. By analyzing action effects on policies and value functions across cooperative and competitive MARL tasks, our method identifies which agent behaviors are beneficial to team success, either by fostering deterministic decisions or by preserving flexibility for future action choices, while also revealing the extent to which agents adopt similar or diverse strategies. Our proposed method offers novel insights into cooperation dynamics and enhances explainability in MARL systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。