用大模型代理模拟学术引用网络,揭示真实世界中的学术行为规律。
Leveraging LLM-based agents for social science research: insights from citation network simulations
- 构建基于大模型代理的引文网络仿真框架CiteAgent
- 成功复现幂律分布、引文扭曲等真实引文网络特征
- 提出两种大模型驱动的社会科学研究范式,适用于学术行为分析
大型语言模型(LLMs)通过海量网络数据预训练,展现出模拟人类行为逻辑与模式的潜力。然而,其在社会模拟中的边界尚不清晰。为此,我们提出CiteAgent框架,利用基于大模型的代理生成引文网络,成功再现了现实引文网络中的主流现象,包括幂律分布、引文扭曲和直径缩小。在此基础上,我们建立了两种大模型驱动的社会科学研究范式:LLM-SE(大模型调查实验)与LLM-LE(大模型实验室实验),支持对引文网络现象的严谨分析,验证并挑战现有理论。同时,通过理想化社会实验拓展了传统科学学研究范畴,仿真结果为真实学术环境提供了宝贵洞见。本工作展示了大模型在推动科学学研究方面的潜力。
原文摘要 · Abstract (English)
The emergence of Large Language Models (LLMs) demonstrates their potential to encapsulate the logic and patterns inherent in human behavior simulation by leveraging extensive web data pre-training. However, the boundaries of LLM capabilities in social simulation remain unclear. To further explore the social attributes of LLMs, we introduce the CiteAgent framework, designed to generate citation networks based on human-behavior simulation with LLM-based agents. CiteAgent successfully captures predominant phenomena in real-world citation networks, including power-law distribution, citational distortion, and shrinking diameter. Building on this realistic simulation, we establish two LLM-based research paradigms in social science: LLM-SE (LLM-based Survey Experiment) and LLM-LE (LLM-based Laboratory Experiment). These paradigms facilitate rigorous analyses of citation network phenomena, allowing us to validate and challenge existing theories. Additionally, we extend the research scope of traditional science of science studies through idealized social experiments, with the simulation experiment results providing valuable insights for real-world academic environments. Our work demonstrates the potential of LLMs for advancing science of science research in social science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。