提出可解释的协作概念学习框架,提升多智能体强化学习透明度与性能。
Concept Learning for Cooperative Multi-Agent Reinforcement Learning
- 用人类可理解的协作概念替代黑箱表示,实现值函数分解的可解释性
- 在星际争霸微操和基于层级的觅食任务中优于现有方法,性能更优
- 支持测试时概念干预,可检测协作模式偏见与虚假关联
尽管神经网络在多智能体强化学习领域取得显著进展,但仍面临透明度与互操作性不足的问题。现有方法依赖黑箱网络,其合作机制难以理解。本文提出一种基于概念瓶颈的可解释值分解框架,通过条件化信用分配至人类可理解的合作概念层面,增强可信度。提出新的基于值的方法CMQ,将每个协作概念表示为监督向量,突破性能与可解释性的权衡。相比传统端到端模型,该方法利用全局状态嵌入对个体动作价值进行条件化,提升协作表征能力。在星际争霸II微操挑战和基于层级的觅食(LBF)任务上的实验表明,CMQ性能优于当前最优方法,且能捕捉有意义的协作模式。此外,支持测试时的概念干预,可识别协作模式中的潜在偏见及影响合作的虚假特征。
原文摘要 · Abstract (English)
Despite substantial progress in applying neural networks (NN) to multi-agent reinforcement learning (MARL) areas, they still largely suffer from a lack of transparency and interoperability. However, its implicit cooperative mechanism is not yet fully understood due to black-box networks. In this work, we study an interpretable value decomposition framework via concept bottleneck models, which promote trustworthiness by conditioning credit assignment on an intermediate level of human-like cooperation concepts. To address this problem, we propose a novel value-based method, named Concepts learning for Multi-agent Q-learning (CMQ), that goes beyond the current performance-vs-interpretability trade-off by learning interpretable cooperation concepts. CMQ represents each cooperation concept as a supervised vector, as opposed to existing models where the information flowing through their end-to-end mechanism is concept-agnostic. Intuitively, using individual action value conditioning on global state embeddings to represent each concept allows for extra cooperation representation capacity. Empirical evaluations on the StarCraft II micromanagement challenge and level-based foraging (LBF) show that CMQ achieves superior performance compared with the state-of-the-art counterparts. The results also demonstrate that CMQ provides more cooperation concept representation capturing meaningful cooperation modes, and supports test-time concept interventions for detecting potential biases of cooperation mode and identifying spurious artifacts that impact cooperation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。