用大模型生成虚拟人格虽高效,但存在系统性偏差,影响预测准确性。
LLM Generated Persona is a Promise with a Catch

- 基于大模型的虚拟人格生成方法缺乏严谨设计,依赖经验性技巧。
- 实验显示生成人格在总统选举和民意调查中与真实数据偏差显著。
- 适合关注生成内容可信度的研究者与政策模拟应用者。
利用大语言模型(LLMs)模拟人类行为受到广泛关注,尤其体现在通过虚拟人格近似个体特征方面。基于人格的模拟有望改变依赖群体反馈的领域,如社会学、经济分析、营销研究和商业运营。传统获取真实人格数据的方法成本高昂、物流复杂且受隐私限制,常无法捕捉多维属性,尤其是主观特质。因此,以大模型生成合成人格成为一种可扩展、低成本的替代方案。然而,当前方法依赖非系统性的启发式生成技术,缺乏方法论严谨性和模拟精度,导致下游任务中出现系统性偏差。通过大规模实验,包括美国总统选举预测和美国人口普遍意见调查,我们发现这些偏差可能导致与现实结果的显著偏离。研究强调需建立人格生成的科学体系,并提出方法创新、组织支持与实证基础的必要性。为促进该领域发展,我们已开源约一百万个生成的人格数据,公开访问地址为 https://huggingface.co/datasets/Tianyi-Lab/Personas。
原文摘要 · Abstract (English)
The use of large language models (LLMs) to simulate human behavior has gained significant attention, particularly through personas that approximate individual characteristics. Persona-based simulations hold promise for transforming disciplines that rely on population-level feedback, including social science, economic analysis, marketing research, and business operations. Traditional methods to collect realistic persona data face significant challenges. They are prohibitively expensive and logistically challenging due to privacy constraints, and often fail to capture multi-dimensional attributes, particularly subjective qualities. Consequently, synthetic persona generation with LLMs offers a scalable, cost-effective alternative. However, current approaches rely on ad hoc and heuristic generation techniques that do not guarantee methodological rigor or simulation precision, resulting in systematic biases in downstream tasks. Through extensive large-scale experiments including presidential election forecasts and general opinion surveys of the U.S. population, we reveal that these biases can lead to significant deviations from real-world outcomes. Our findings underscore the need to develop a rigorous science of persona generation and outline the methodological innovations, organizational and institutional support, and empirical foundations required to enhance the reliability and scalability of LLM-driven persona simulations. To support further research and development in this area, we have open-sourced approximately one million generated personas, available for public access and analysis at https://huggingface.co/datasets/Tianyi-Lab/Personas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。