arXiv:2509.18088cs.MAcs.LG2025-09被引 3

提出分层强化与集体学习框架,解决动态多智能体系统的协同优化难题。

Strategic Coordination for Evolving Multi-agent Systems: A Hierarchical Reinforcement and Collective Learning Approach

  • 分层架构:高层用MARL生成策略,降低动作空间;低层集体学习实现去中心化协调。
  • 在真实城市场景中,性能优于单一MARL或集体学习方法,提升可扩展性与适应性。
  • 适合需自主决策且应对突发变化的复杂多智能体系统,如智慧城市能源管理。

去中心化的组合优化在动态多智能体系统中面临重大挑战,要求智能体在长期决策与短期集体最优之间权衡,同时在未预期变化下保持交互自主性。强化学习可通过动态规划建模序列决策以预测环境变化,但将多智能体强化学习(MARL)应用于此类问题仍具挑战,受限于联合状态-动作空间指数增长、高通信开销及集中训练带来的隐私问题。为此,本文提出分层强化与集体学习(HRCL)新方法,结合MARL与基于分层框架的去中心化集体学习。高层通过MARL生成策略,对可行计划进行分组以缩减动作空间,并约束智能体行为以实现帕累托最优;低层集体学习层则确保智能体间高效、去中心化协调,通信开销极小。在合成场景与真实世界智慧城市应用模型(包括能源自管理与无人机群感知)中的大量实验表明,相比独立使用MARL或集体学习,HRCL显著提升性能、可扩展性与适应性,达成共赢协同方案。

原文摘要 · Abstract (English)

Decentralized combinatorial optimization in evolving multi-agent systems poses significant challenges, requiring agents to balance long-term decision-making, short-term optimized collective outcomes, while preserving autonomy of interactive agents under unanticipated changes. Reinforcement learning offers a way to model sequential decision-making through dynamic programming to anticipate future environmental changes. However, applying multi-agent reinforcement learning (MARL) to decentralized combinatorial optimization problems remains an open challenge due to the exponential growth of the joint state-action space, high communication overhead, and privacy concerns in centralized training. To address these limitations, this paper proposes Hierarchical Reinforcement and Collective Learning (HRCL), a novel approach that leverages both MARL and decentralized collective learning based on a hierarchical framework. Agents take high-level strategies using MARL to group possible plans for action space reduction and constrain the agent behavior for Pareto optimality. Meanwhile, the low-level collective learning layer ensures efficient and decentralized coordinated decisions among agents with minimal communication. Extensive experiments in a synthetic scenario and real-world smart city application models, including energy self-management and drone swarm sensing, demonstrate that HRCL significantly improves performance, scalability, and adaptability compared to the standalone MARL and collective learning approaches, achieving a win-win synthesis solution.

多智能体强化学习协同优化分层学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。