2058人问卷数据集助力构建高精度人类数字孪生模型
Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions
- 基于2058人、500题跨波次问卷构建个体行为数据集
- 涵盖心理、经济、认知等多维度指标,支持行为预测与验证
- 适合研究数字孪生、社会科学研究及个性化建模的团队使用
基于大语言模型的人类数字孪生模拟在人工智能、社会科学和数字实验领域具有广阔前景,但受限于高质量、大规模且公开的个体级数据稀缺。为弥补这一空白,我们构建了一个大规模公开数据集,旨在全面刻画个体行为。通过对美国代表性样本共2,058名参与者进行四轮调查(每人平均耗时2.42小时),覆盖总计500个问题,内容包括人口统计、心理、经济、人格、认知测量,以及行为经济学实验复现与定价调查。最后一轮重复前期任务,用于建立重测信度基准。初步分析表明数据质量高,具备构建精准个体与群体行为预测数字孪生的潜力。通过公开完整数据集,我们期望为基于大语言模型的人物模拟提供重要测试平台。此外,其独特广度与规模也支持广泛的社会科学研究,如跨构念相关性分析与异质处理效应研究。
原文摘要 · Abstract (English)
LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-quality ground truth limits both the development and validation of digital twin methodologies. To address this gap, we introduce a large-scale, public dataset designed to capture a rich and holistic view of individual human behavior. We survey a representative sample of $N = 2,058$ participants (average 2.42 hours per person) in the US across four waves with 500 questions in total, covering a comprehensive battery of demographic, psychological, economic, personality, and cognitive measures, as well as replications of behavioral economics experiments and a pricing survey. The final wave repeats tasks from earlier waves to establish a test-retest accuracy baseline. Initial analyses suggest the data are of high quality and show promise for constructing digital twins that predict human behavior well at the individual and aggregate levels. By making the full dataset publicly available, we aim to establish a valuable testbed for the development and benchmarking of LLM-based persona simulations. Beyond LLM applications, due to its unique breadth and scale the dataset also enables broad social science research, including studies of cross-construct correlations and heterogeneous treatment effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。