用100个智能体模拟十年社会生活,让大模型学会人类社交智慧。
Agentopia: Long-Term Life Simulation and Learning in Agent Societies

- 构建100个智能体的长期社会模拟系统,持续10年自主发展
- 通过生命奖励机制训练大模型,角色幸福感提升,角色扮演能力提高15.6%
- 适合对社会行为建模、通用智能体训练感兴趣的读者
人类从社会生活中学习。用大语言模型驱动的智能体模拟这一过程是一个有前景的研究方向,核心问题是:大模型能否通过这种模拟的社会经验来更好地理解并复现人类行为?然而,以往的智能体社会模拟通常仅持续数天,限制了社会互动深度和长期成长。本文研究智能体社会中的长期生活模拟与大模型学习,目标有两个:(1) 探究终身模拟中涌现的社会行为;(2) 通过多年模拟社会经验,发展大模型的人类化能力,尤其是社会智力。我们提出 Agentopia,一个面向多智能体社会的长期生活模拟综合框架,其中100个智能体在10个模拟年内自主追求个人成长、建立社会关系、满足需求与目标。我们定义‘生命奖励’以反映人类福祉,并利用该奖励通过拒绝采样训练大模型。大量实验表明,智能体展现出丰富的涌现社会行为。此外,生命奖励训练有效提升了底层大模型,使模拟中的智能体福祉提升,并在下游角色扮演基准上取得+15.6%的性能提升。
原文摘要 · Abstract (English)
Humans learn from social life. Simulating this process with LLM-powered agents represents a promising research direction, raising a natural question: whether LLMs can learn from such simulated social experience to better understand and replicate human behavior. However, prior agent society simulations typically operate at the scale of days, limiting the depth of social interactions and long-term growth. In this paper, we study long-term life simulation and LLM learning in agent societies, with two goals: (1) investigating social behaviors that emerge from life-long simulation, and (2) developing anthropomorphic capabilities in LLMs, particularly intelligence in social life, through years of simulated social experience. Specifically, we present Agentopia, a comprehensive framework for long-term life simulation in multi-agent societies, where 100 agents autonomously pursue personal growth, develop social relationships, and fulfill their needs and goals over 10 simulated years. We define life reward to mirror human well-being, and leverage this reward to train LLMs via rejection sampling. Extensive experiments show that agents exhibit rich emergent social behaviors. Furthermore, life reward training effectively enhances the underlying LLM, which leads to improved agent well-being in simulation, and generalizes to downstream role-playing benchmarks with +15.6% improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。