arXiv:2509.16709cs.LG2025-09被引 1

用超网络让多智能体协同控制高维系统,效果更好且更省资源。

HypeMARL: Multi-Agent Reinforcement Learning For High-Dimensional, Parametric, and Distributed Systems

  • 用超网络+位置编码,让智能体根据参数和位置自适应策略
  • 在密度与流动控制任务中性能超越现有去中心化方法
  • 只需极少调参,且可减少约90%的环境交互次数

深度强化学习正成为控制由偏微分方程(PDE)描述的复杂动态系统的有前景策略。针对状态与控制变量高维、分布式的难题,多智能体强化学习(MARL)通过去中心化训练与执行,可缓解维度灾难。然而,当智能体需表现集体非局部行为以最大化奖励时,局部性原则可能成为瓶颈,这在典型PDE约束最优控制问题中常见。本文提出HypeMARL:一种专为高维、参数化、分布式系统设计的去中心化MARL算法。该方法利用超网络,基于系统参数和智能体相对位置(通过正弦位置编码表示)有效参数化智能体策略与价值函数。在密度与流场控制等挑战性任务中验证表明,HypeMARL(i)能通过智能体集体行为实现高效控制,优于当前最先进的去中心化MARL;(ii)可有效处理参数依赖;(iii)所需超参数调整极少;(iv)其模型基扩展MB-HypeMARL借助计算高效的深度学习代理模型局部近似系统动力学,使环境交互次数减少约10倍,同时政策性能损失极小。

原文摘要 · Abstract (English)

Deep reinforcement learning has recently emerged as a promising feedback control strategy for complex dynamical systems governed by partial differential equations (PDEs). When dealing with distributed, high-dimensional problems in state and control variables, multi-agent reinforcement learning (MARL) has been proposed as a scalable approach for breaking the curse of dimensionality. In particular, through decentralized training and execution, multiple agents cooperate to steer the system towards a target configuration, relying solely on local state and reward information. However, the principle of locality may become a limiting factor whenever a collective, nonlocal behavior of the agents is crucial to maximize the reward function, as typically happens in PDE-constrained optimal control problems. In this work, we propose HypeMARL: a decentralized MARL algorithm tailored to the control of high-dimensional, parametric, and distributed systems. HypeMARL employs hypernetworks to effectively parametrize the agents' policies and value functions with respect to the system parameters and the agents' relative positions, encoded by sinusoidal positional encoding. Through the application on challenging control problems, such as density and flow control, we show that HypeMARL (i) can effectively control systems through a collective behavior of the agents, outperforming state-of-the-art decentralized MARL, (ii) can efficiently deal with parametric dependencies, (iii) requires minimal hyperparameter tuning and (iv) can reduce the amount of expensive environment interactions by a factor of ~10 thanks to its model-based extension, MB-HypeMARL, which relies on computationally efficient deep learning-based surrogate models approximating the dynamics locally, with minimal deterioration of the policy performance.

多智能体强化学习高维控制超网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。