用大模型模拟社交网络行为前,必须验证其真实度。
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
- 构建社交网络仿真框架,测试大模型模仿用户沟通
- 实证显示不同方法在英德语境下表现差异显著
- 强调必须在实际场景中验证模拟真实性
大型语言模型(LLMs)模拟人类行为的能力引发了大量计算社会科学研究,研究者普遍假设可用AI代理替代真人进行实证分析。然而,关于该假设是否成立仍存在分歧,亟需厘清实验设计的差异。本文聚焦于利用LLMs模拟社交网络用户行为,尤其关注通信模式的再现。我们首先提出一个形式化社交网络仿真框架,随后重点研究用户沟通的模仿。在X平台(原Twitter)上,针对英语和德语用户,我们实证测试了多种模仿策略。结果表明,社交仿真必须通过在拟合时所用场景下的经验真实性来验证。本文主张,在使用生成式代理进行社会仿真时应更加严谨。
原文摘要 · Abstract (English)
The ability of Large Language Models (LLMs) to mimic human behavior triggered a plethora of computational social science research, assuming that empirical studies of humans can be conducted with AI agents instead. Since there have been conflicting research findings on whether and when this hypothesis holds, there is a need to better understand the differences in their experimental designs. We focus on replicating the behavior of social network users with the use of LLMs for the analysis of communication on social networks. First, we provide a formal framework for the simulation of social networks, before focusing on the sub-task of imitating user communication. We empirically test different approaches to imitate user behavior on X in English and German. Our findings suggest that social simulations should be validated by their empirical realism measured in the setting in which the simulation components were fitted. With this paper, we argue for more rigor when applying generative-agent-based modeling for social simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。