让大模型智能体通过互相讨论自动进化,无需外部奖励。
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- 用智能体间互动生成内在奖励信号,驱动自我改进。
- 在多数任务中超越未训练智能体,达到当前最佳表现。
- 适合研究自主学习与多智能体协同的学者参考。
自进化是使大语言模型(LLM)智能体在预训练后持续提升能力的核心课题。近期研究从无强化学习转向基于强化学习的方法。现有方法或依赖密集外部奖励信号,或从LLM自身提取内在奖励信号。然而,这些方法偏离了人类智能中通过相互讨论与协作实现学习的机制。本文提出协同演化多智能体系统(CoMAS),一种使智能体通过交互自主进化的新型框架。CoMAS从丰富的讨论动态中生成内在奖励,利用LLM作为评判者构建奖励信号,并通过强化学习优化各智能体策略,从而实现去中心化、可扩展的协同演化。实验表明,CoMAS在多数评估设置下持续优于未训练智能体,达到当前最优性能。消融实验证实互动奖励信号的必要性,并揭示随着智能体数量与多样性增加,其表现出良好可扩展性。这些发现确立了CoMAS作为LLM智能体自进化的新范式。
原文摘要 · Abstract (English)
Self-evolution is a central research topic in enabling large language model (LLM)-based agents to continually improve their capabilities after pretraining. Recent research has witnessed a transition from reinforcement learning (RL)-free to RL-based methods. Current RL-based methods either rely on dense external reward signals or extract intrinsic reward signals from LLMs themselves. However, these approaches diverge from the self-evolution mechanisms observed in human intelligence, where individuals learn and improve through mutual discussion and collaboration. In this work, we introduce Co-Evolving Multi-Agent Systems (CoMAS), a novel framework that enables agents to improve autonomously by learning from inter-agent interactions without external supervision. CoMAS generates intrinsic rewards from rich discussion dynamics, employs an LLM-as-a-judge mechanism to formulate these rewards, and optimizes each agent's policy through RL, thereby enabling decentralized and scalable co-evolution. Experimental results demonstrate that CoMAS consistently outperforms untrained agents and achieves state-of-the-art performance across most evaluation settings. Ablation studies confirm the necessity of interaction-based reward signals and reveal promising scalability as the number and diversity of agents increase. These findings establish CoMAS as a novel and effective paradigm for self-evolution in LLM-based agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。