arXiv:2502.12605cs.MAcs.LG2025-02

用超网络统一优化可控与不可控智能体的配置和策略,提升多智能体系统效率。

Hypernetwork-based approach for optimal composition design in partially controlled multi-agent systems

  • 通过统一超网络生成各类配置下的智能体策略,避免重复训练。
  • 在纽约出租车数据上实现订单响应率与服务需求量显著提升。
  • 适合需要高效设计多智能体系统架构的研究者与工程师。

部分可控多智能体系统(PCMAS)由系统设计者控制的可控智能体和自主运行的不可控智能体构成。本文研究了PCMAS中的最优组合设计问题,即系统设计者需确定最优的可控智能体数量与策略,而不可控智能体需找到其最优响应策略。该双层优化问题计算复杂,因需在不同配置下反复求解多智能体强化学习问题。为此,我们提出一种基于超网络的新框架,联合优化系统组合与智能体策略。与传统方法为每种配置单独训练策略网络不同,本框架通过统一超网络生成可控与不可控智能体的策略,实现相似配置间的高效信息共享,降低计算开销。进一步结合奖励参数优化与平均动作网络,提升性能。基于真实纽约市出租车数据实验表明,该框架在逼近均衡策略方面优于现有方法,关键指标如订单响应率和服务需求量显著提升,验证了可控智能体在决策优化中的实际价值。

原文摘要 · Abstract (English)

Partially Controlled Multi-Agent Systems (PCMAS) are comprised of controllable agents, managed by a system designer, and uncontrollable agents, operating autonomously. This study addresses an optimal composition design problem in PCMAS, which involves the system designer's problem, determining the optimal number and policies of controllable agents, and the uncontrollable agents' problem, identifying their best-response policies. Solving this bi-level optimization problem is computationally intensive, as it requires repeatedly solving multi-agent reinforcement learning problems under various compositions for both types of agents. To address these challenges, we propose a novel hypernetwork-based framework that jointly optimizes the system's composition and agent policies. Unlike traditional methods that train separate policy networks for each composition, the proposed framework generates policies for both controllable and uncontrollable agents through a unified hypernetwork. This approach enables efficient information sharing across similar configurations, thereby reducing computational overhead. Additional improvements are achieved by incorporating reward parameter optimization and mean action networks. Using real-world New York City taxi data, we demonstrate that our framework outperforms existing methods in approximating equilibrium policies. Our experimental results show significant improvements in key performance metrics, such as order response rate and served demand, highlighting the practical utility of controlling agents and their potential to enhance decision-making in PCMAS.

多智能体超网络强化学习系统优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。