开放环境中多智能体强化学习的贡献分配难题
Challenges in Credit Assignment for Multi-Agent Reinforcement Learning in Open Agent Systems
- 区分三类开放性:智能体进出、任务演化、能力变化
- 开放导致信用误分配,损失函数不稳,性能显著下降
- 适合研究动态系统与多智能体协作的学者参考
在快速发展的多智能体强化学习(MARL)领域,理解开放系统的动态特性至关重要。开放性指系统中智能体数量、任务类型及智能体能力随时间动态变化。具体包括三类开放性:智能体开放性(可随时进入或离开)、任务开放性(新任务出现或旧任务演化消失)以及类型开放性(智能体能力与行为随时间改变)。本文从概念与实证角度分析开放性与信用分配问题(CAP)之间的相互作用。CAP旨在确定个体智能体对整体系统表现的贡献,但在开放环境中其复杂度急剧上升。传统信用分配方法通常假设智能体群体静态、任务固定且类型不变,难以适应开放系统。我们首先进行概念分析,提出新的开放性子类别,说明智能体更替或任务取消如何破坏环境平稳性与团队结构固定的假设。随后,在开放环境中对典型的时间与结构算法进行实证研究,结果表明开放性直接引发信用误分配,表现为损失函数不稳定和性能显著下降。
原文摘要 · Abstract (English)
In the rapidly evolving field of multi-agent reinforcement learning (MARL), understanding the dynamics of open systems is crucial. Openness in MARL refers to the dynam-ic nature of agent populations, tasks, and agent types with-in a system. Specifically, there are three types of openness as reported in (Eck et al. 2023) [2]: agent openness, where agents can enter or leave the system at any time; task openness, where new tasks emerge, and existing ones evolve or disappear; and type openness, where the capabil-ities and behaviors of agents change over time. This report provides a conceptual and empirical review, focusing on the interplay between openness and the credit assignment problem (CAP). CAP involves determining the contribution of individual agents to the overall system performance, a task that becomes increasingly complex in open environ-ments. Traditional credit assignment (CA) methods often assume static agent populations, fixed and pre-defined tasks, and stationary types, making them inadequate for open systems. We first conduct a conceptual analysis, in-troducing new sub-categories of openness to detail how events like agent turnover or task cancellation break the assumptions of environmental stationarity and fixed team composition that underpin existing CAP methods. We then present an empirical study using representative temporal and structural algorithms in an open environment. The results demonstrate that openness directly causes credit misattribution, evidenced by unstable loss functions and significant performance degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。