验证大模型在协作任务中可持续合作能力,拓展了适用场景与语言
Reproducibility Study of "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents"
- 复现并扩展原框架,测试多模型在资源共享中的协作表现
- 大模型(如GPT-4-turbo)无论是否启用通用化原则均能持续合作
- 适用于新模型、多语言及异构环境,对智能体系统设计有指导意义
本研究评估并扩展了Piatti等人提出的GovSim模拟框架,该框架用于评估大语言模型(LLMs)在资源分享场景中的协作决策能力。通过复现关键实验,验证了大型模型(如GPT-4-turbo)相较于小型模型的协作性能优势。结果表明,大型模型即使不使用通用化原则也能实现可持续合作,而小型模型则依赖该原则。此外,本研究引入多个扩展:测试了DeepSeek-V3和GPT-4o-mini等新模型,验证了协作行为在不同架构与规模下的泛化性;构建异构多智能体环境,研究日语指令场景,以及探索“逆向环境”——智能体需合作缓解有害资源分配。结果证实该基准可推广至新模型、新场景与多语言,且高绩效模型能引导低绩效模型模仿其行为,这对智能体系统优化具有重要启示。
原文摘要 · Abstract (English)
This study evaluates and extends the findings made by Piatti et al., who introduced GovSim, a simulation framework designed to assess the cooperative decision-making capabilities of large language models (LLMs) in resource-sharing scenarios. By replicating key experiments, we validate claims regarding the performance of large models, such as GPT-4-turbo, compared to smaller models. The impact of the universalization principle is also examined, with results showing that large models can achieve sustainable cooperation, with or without the principle, while smaller models fail without it. In addition, we provide multiple extensions to explore the applicability of the framework to new settings. We evaluate additional models, such as DeepSeek-V3 and GPT-4o-mini, to test whether cooperative behavior generalizes across different architectures and model sizes. Furthermore, we introduce new settings: we create a heterogeneous multi-agent environment, study a scenario using Japanese instructions, and explore an "inverse environment" where agents must cooperate to mitigate harmful resource distributions. Our results confirm that the benchmark can be applied to new models, scenarios, and languages, offering valuable insights into the adaptability of LLMs in complex cooperative tasks. Moreover, the experiment involving heterogeneous multi-agent systems demonstrates that high-performing models can influence lower-performing ones to adopt similar behaviors. This finding has significant implications for other agent-based applications, potentially enabling more efficient use of computational resources and contributing to the development of more effective cooperative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。