arXiv:2510.22477cs.MAcs.AI2025-10被引 1

用序列强化学习优化通信效率,让多智能体更省 token

Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization

  • 用 GSPO 算法训练智能体,奖励机制惩罚啰嗦发言
  • 7 个推理任务上性能新高,耗 token 数量仅为现有方法几分之一
  • 能自发产生‘战略性沉默’,适合需低成本通信的系统

为解决‘自由交流’型多智能体系统中高昂的通信成本问题,我们提出 Agent-GSPO 框架,通过序列级强化学习直接优化令牌经济。该框架利用稳定且内存高效的群体序列策略优化(GSPO)算法,在显式惩罚冗余表达的通信感知奖励下训练智能体。在七个推理基准测试中,Agent-GSPO 不仅达到新的最优性能,且消耗的令牌数量仅为现有方法的极小部分。通过催生如‘战略性沉默’等涌现策略,本方法为构建可扩展、经济可行的多智能体系统提供了实用蓝图。

原文摘要 · Abstract (English)

To combat the prohibitive communication costs of ``free-for-all" multi-agent systems (MAS), we introduce \textbf{Agent-GSPO}, a framework that directly optimizes for token economy using sequence-level reinforcement learning. Agent-GSPO leverages the stable and memory-efficient Group Sequence Policy Optimization (GSPO) algorithm to train agents on a communication-aware reward that explicitly penalizes verbosity. Across seven reasoning benchmarks, Agent-GSPO not only achieves new state-of-the-art performance but does so with a fraction of the token consumption of existing methods. By fostering emergent strategies like ``strategic silence," our approach provides a practical blueprint for developing scalable and economically viable multi-agent systems.

多智能体通信效率强化学习令牌经济

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。