提出直接可解释性方法,让训练好的多智能体模型自己说明行为逻辑。
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
- 不修改模型结构,从已训练模型生成行为解释
- 涵盖归因反传、激活修补等六种现代解释技术
- 适合研究团队协作与智能体偏差的学者
多智能体深度强化学习在机器人或游戏等复杂问题中表现高效,但多数训练模型难以解释。虽然内在可解释模型虽具潜力,但在处理复杂任务或多智能体动态时存在扩展性和灵活性不足的问题。本文主张采用直接可解释性,即从已训练模型直接生成事后解释,作为灵活且可扩展的替代方案,无需改动模型架构即可揭示智能体行为、涌现现象及偏见。我们探讨了现代方法,包括相关性反向传播、知识编辑、模型调控、激活修补、稀疏自编码器和电路发现,展示其在单智能体、多智能体及训练过程中的适用性。通过解决多智能体强化学习的可解释性问题,本文提出若干方向,旨在推进团队识别、群体协调和样本效率等活跃课题。
原文摘要 · Abstract (English)
Multi-Agent Deep Reinforcement Learning (MADRL) was proven efficient in solving complex problems in robotics or games, yet most of the trained models are hard to interpret. While learning intrinsically interpretable models remains a prominent approach, its scalability and flexibility are limited in handling complex tasks or multi-agent dynamics. This paper advocates for direct interpretability, generating post hoc explanations directly from trained models, as a versatile and scalable alternative, offering insights into agents' behaviour, emergent phenomena, and biases without altering models' architectures. We explore modern methods, including relevance backpropagation, knowledge edition, model steering, activation patching, sparse autoencoders and circuit discovery, to highlight their applicability to single-agent, multi-agent, and training process challenges. By addressing MADRL interpretability, we propose directions aiming to advance active topics such as team identification, swarm coordination and sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。