arXiv:2510.26389cs.LGcs.MA2025-10NeurIPS被引 8

动态调整上下文长度,用低频信息过滤冗余,提升多智能体强化学习效率。

Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement Learning

  • 通过时序梯度分析动态优化上下文长度
  • 在多个长依赖任务上达到当前最优性能
  • 适合处理具有长期依赖的复杂多智能体场景

近期深度多智能体强化学习(MARL)在解决长时依赖和非马尔可夫环境等挑战性任务上表现优异,部分归功于对大固定上下文长度的依赖。然而,固定长上下文可能导致探索效率低下和信息冗余。本文提出一种新型MARL框架,实现自适应、高效的上下文信息获取。具体地,设计一个中心智能体,通过时序梯度分析动态优化上下文长度,提升探索能力,促进收敛至全局最优。此外,为增强上下文长度的自适应优化能力,提出一种高效输入表示方法,利用基于傅里叶的低频截断技术,提取各分散智能体间的全局时序趋势,提供有效且高效的环境表征。大量实验表明,该方法在包含PettingZoo、MiniGrid、Google Research Football (GRF) 和 StarCraft Multi-Agent Challenge v2 (SMACv2) 的长时依赖任务上均达到当前最优(SOTA)性能。

原文摘要 · Abstract (English)

Recently, deep multi-agent reinforcement learning (MARL) has demonstrated promising performance for solving challenging tasks, such as long-term dependencies and non-Markovian environments. Its success is partly attributed to conditioning policies on large fixed context length. However, such large fixed context lengths may lead to limited exploration efficiency and redundant information. In this paper, we propose a novel MARL framework to obtain adaptive and effective contextual information. Specifically, we design a central agent that dynamically optimizes context length via temporal gradient analysis, enhancing exploration to facilitate convergence to global optima in MARL. Furthermore, to enhance the adaptive optimization capability of the context length, we present an efficient input representation for the central agent, which effectively filters redundant information. By leveraging a Fourier-based low-frequency truncation method, we extract global temporal trends across decentralized agents, providing an effective and efficient representation of the MARL environment. Extensive experiments demonstrate that the proposed method achieves state-of-the-art (SOTA) performance on long-term dependency tasks, including PettingZoo, MiniGrid, Google Research Football (GRF), and StarCraft Multi-Agent Challenge v2 (SMACv2).

多智能体强化学习上下文优化低频滤波

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。