提出靶向干预机制,让单个智能体高效引导多智能体协作。
A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
- 用多智能体影响图分析协作机制,识别可干预节点。
- 通过因果推断实现仅干预单一智能体,提升全局协同效率。
- 适用于大规模多智能体系统,无需全局控制信号。
在大规模多智能体强化学习(MARL)中,依赖人类全局指导难以实现,而现有外部协调机制多依赖经验设计,缺乏理论支撑。本文引入多智能体影响图(MAIDs)作为图形化框架,提出一种新交互范式——靶向干预,仅对单个目标智能体施加干预,缓解全局引导难题。基于MAIDs的因果结构,设计预策略干预(PSI)技术,通过最大化因果效应实现主任务目标与附加目标的联合优化。此外,利用MAIDs的捆绑相关图分析,可判断特定学习范式在给定交互设计下的可行性。实验验证了靶向干预的有效性及相关图分析的准确性。
原文摘要 · Abstract (English)
Steering cooperative multi-agent reinforcement learning (MARL) towards desired outcomes is challenging, particularly when the global guidance from a human on the whole multi-agent system is impractical in a large-scale MARL. On the other hand, designing external mechanisms (e.g., intrinsic rewards and human feedback) to coordinate agents mostly relies on empirical studies, lacking a easy-to-use research tool. In this work, we employ multi-agent influence diagrams (MAIDs) as a graphical framework to address the above issues. First, we introduce the concept of MARL interaction paradigms (orthogonal to MARL learning paradigms), using MAIDs to analyze and visualize both unguided self-organization and global guidance mechanisms in MARL. Then, we design a new MARL interaction paradigm, referred to as the targeted intervention paradigm that is applied to only a single targeted agent, so the problem of global guidance can be mitigated. In implementation, we introduce a causal inference technique, referred to as Pre-Strategy Intervention (PSI), to realize the targeted intervention paradigm. Since MAIDs can be regarded as a special class of causal diagrams, a composite desired outcome that integrates the primary task goal and an additional desired outcome can be achieved by maximizing the corresponding causal effect through the PSI. Moreover, the bundled relevance graph analysis of MAIDs provides a tool to identify whether an MARL learning paradigm is workable under the design of an MARL interaction paradigm. In experiments, we demonstrate the effectiveness of our proposed targeted intervention, and verify the result of relevance graph analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。