用多重人格模型模拟人群行为,让大模型更真实地反映个体差异。
Mixture-of-Personas Language Models for Population Simulation
- 基于概率混合机制,为每个虚拟人分配独特人格与行为样本。
- 在合成数据生成中,对齐度和多样性均优于现有方法。
- 无需微调,可跨不同大模型使用,适合社会科学研究者。
大型语言模型(LLMs)在人类行为模拟等领域的应用日益广泛,可补充社会科学与机器学习中的数据。然而,预训练的LLM常因个体与群体间固有的行为差异,难以捕捉目标人群的多样性。为此,我们提出「多重人格」(Mixture of Personas, MoP)——一种概率化提示方法,使模型输出更贴合目标人群。MoP是一种上下文混合模型,每个成分由一个具有特定人格的LM代理及其代表子群体行为的示例组成。通过根据学习到的混合权重随机选取人格与示例,激发大模型生成多样化响应。该方法无需模型微调,具备灵活性与跨模型可迁移性。实验表明,在合成数据生成任务中,MoP在对齐度与多样性指标上均优于对比方法。
原文摘要 · Abstract (English)
Advances in Large Language Models (LLMs) paved the way for their emerging applications in various domains, such as human behavior simulations, where LLMs could augment human-generated data in social science research and machine learning model training. However, pretrained LLMs often fail to capture the behavioral diversity of target populations due to the inherent variability across individuals and groups. To address this, we propose \textit{Mixture of Personas} (MoP), a \textit{probabilistic} prompting method that aligns the LLM responses with the target population. MoP is a contextual mixture model, where each component is an LM agent characterized by a persona and an exemplar representing subpopulation behaviors. The persona and exemplar are randomly chosen according to the learned mixing weights to elicit diverse LLM responses during simulation. MoP is flexible, requires no model finetuning, and is transferable across base models. Experiments for synthetic data generation show that MoP outperforms competing methods in alignment and diversity metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。