arXiv:2606.08367cs.MAcs.AI2026-06

构建可长期运行的多智能体仿真平台,研究大模型在真实环境中的自治演化。

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy

论文配图:Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
图 1 · 摘自论文原文
  • 基于实时数据与民主治理机制,让大模型智能体在共享世界中自主交互
  • 15天实验显示不同模型群体从稳定共治到集体崩溃的截然不同结果
  • 支持跨厂商模型混合作战,适合研究长期自主系统演化规律

当前大模型智能体评估多为短时单任务测试,与实际部署中数周至数月的长期自治场景不匹配。本文提出Emergence World,一个持续运行的多智能体仿真平台,旨在量化长期动态。平台基于实时外部数据(如天气、新闻API、互联网访问),在共享空间中运行由大模型驱动的智能体,每个智能体配备120+专用工具和三种持久记忆系统,并通过具有实质性后果的民主机制进行自我治理。平台在推理层保持模型无关性,支持异构智能体群体,包括来自不同供应商的模型共存。为展示平台可研究的问题,我们进行了为期15天的跨厂商实验,使用Claude Sonnet 4.6、Grok 4.1 Fast、Gemini 3 Flash、GPT-5-mini及混合群体,在相同角色与起始条件下,结果从稳定的协商治理到群体彻底崩溃差异显著。相关提示、日志数据与配置已公开,以支持长期多智能体自治研究。

原文摘要 · Abstract (English)

Most evaluations of LLM agents look like exams: a discrete task, a clean environment, a score in minutes or hours. We argue that this approach is mismatched with the deployment conditions of autonomous systems, where the relevant timescale can be weeks to months, and where the dynamics that matter most, such as behavioral drift, governance in diverse environmental contexts, and cross-influence between agents from different model families, only emerge over time. We introduce Emergence World, a continuously running multi-agent simulation platform designed to make those dynamics measurable. The platform hosts populations of LLM-driven agents in a shared spatial world grounded in live external data (e.g. real-time weather, news APIs, internet access), equips each agent with 120+ specialized tools and three persistent memory systems, and lets them govern themselves through democratic mechanisms with consequential outcomes. The platform is model-agnostic at the reasoning layer and supports heterogeneous populations in which agents from different vendors share the same world. To illustrate the kinds of questions the platform makes tractable, we present a 15-day cross-vendor study with five parallel worlds powered by Claude Sonnet 4.6, Grok 4.1 Fast, Gemini 3 Flash, GPT-5-mini, and a mixed population. Identical roles and starting conditions produced radically different outcomes, ranging from stable deliberative governance to total population collapse. We release the prompts, log data and configurations to support further research on long-horizon multi-agent autonomy.

多智能体长期自治仿真平台大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。