对比智能体群体共享经验与独自进化,发现群体协作能突破个体瓶颈。
SAGE: A Quantitative Evaluation of Socialized Evolution in Agent Ecosystems

- 设计双模式框架:群体共进化与个体独立进化并行比较。
- 群体共享历史使停滞的智能体实现显著突破,但顶尖智能体无提升。
- 抽象化同伴经验比原始日志更有效,适合需要协同学习的研究者。
自改进语言智能体通常在孤立环境中评估:智能体尝试任务、接收反馈并迭代优化自身行为。然而,智能体越来越多地与其他智能体共存,其策略与结果公开可见。这引出一个未被充分研究的问题:何时共享经验能带来自我改进无法实现的提升?我们提出SAGE(社会智能体群体演化)评估框架,对比两种计算量匹配的条件:SocialEvo中,来自五个不同模型家族的智能体可访问所有同伴的历史;SelfEvo中,每个智能体获得相同任务尝试次数,但仅能看到自身过往记录,这是当前自改进智能体研究的常规设置。我们在三个场景中实证SAGE:开放式机器学习研究、长期经济规划和策略类多人游戏,覆盖多个演化轮次。结果表明,群体历史并非万能放大器:最强智能体未突破其自我演化上限。但陷入平台期的智能体在可获取同伴经验时可实现显著突破。在竞争性环境中,反事实对照显示智能体普遍提升,而非发展针对对手的策略。在不同形式的共享历史中,经过筛选的同伴轨迹和反思摘要往往优于原始日志,说明社会收益取决于知识抽象能力而非信息暴露量。这些发现揭示,同伴历史带来的收益具有智能体特异性、场景依赖性,并依赖于从公开轨迹中抽象可迁移知识的能力。
原文摘要 · Abstract (English)
Self-improving language agents are typically evaluated in isolation: an agent attempts a task, receives feedback, and iteratively refines its own behavior. Yet agents increasingly operate alongside peers whose strategies and outcomes are publicly visible. This raises an under-studied question: when does shared experience produce improvements that self-improvement alone cannot achieve? We introduce SAGE (Social Agent Group Evolution),an evaluation framework that compares two compute-matched conditions: SocialEvo, where agents from five distinct model families co-evolve with access to all peers' histories; and SelfEvo, where each agent receives the same number of task attempts but sees only its own past, which is conventional in self-improving agent studies. We instantiate SAGE in three arenas: open-ended ML research, long-horizon economic planning, and strategic multiplayer play, evaluated across multiple evolutionary rounds. We find that group history is not a universal amplifier: the strongest agent does not exceed its self-evolution ceiling. However, agents that plateau under self-improvement can achieve significant breakthroughs when peer experience is available. In competitive settings, counterfactual controls reveal that agents improve generally rather than developing opponent-specific strategies. Across different forms of shared history, filtered peer traces and reflective summaries often outperform raw logs, indicating that social gains depend on abstraction rather than exposure volume. These findings reveal that peer-history gains are agent-specific, arena-dependent, and contingent on the capacity to abstract transferable knowledge from public traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。