arXiv:2512.07890cs.MAcs.AI2025-12被引 5

用生成模型增强大模型,打造更真实数字人群。

CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models

  • 融合预训练大模型与生成模型,提升数字人群多样性。
  • 在众包、投票等场景中,表现接近真实人类数据。
  • 适合社会模拟、营销推荐等需要大规模虚拟用户的研究。

大语言模型(LLMs)的兴起激发了构建基于大模型的数字人群的兴趣,可应用于社会模拟、众包、营销和推荐系统等领域。数字人群能降低招募真人参与者成本,并缓解涉及人类受试者研究的相关伦理问题。然而,现有方法多仅依赖大模型,难以充分捕捉真实人群的准确性和多样性。为此,我们提出 CrowdLLM,通过整合预训练大模型与生成模型,增强数字人群的多样性和真实性。我们从理论上分析了 CrowdLLM 在创建低成本、高代表性、可扩展的数字人群方面的潜力,其质量可媲美真实人群。在多个领域(如众包、投票、用户评分)进行的综合实验与仿真研究证明,CrowdLLM 在准确性和分布保真度方面均表现出色。

原文摘要 · Abstract (English)

The emergence of large language models (LLMs) has sparked much interest in creating LLM-based digital populations that can be applied to many applications such as social simulation, crowdsourcing, marketing, and recommendation systems. A digital population can reduce the cost of recruiting human participants and alleviate many concerns related to human subject study. However, research has found that most of the existing works rely solely on LLMs and could not sufficiently capture the accuracy and diversity of a real human population. To address this limitation, we propose CrowdLLM that integrates pretrained LLMs and generative models to enhance the diversity and fidelity of the digital population. We conduct theoretical analysis of CrowdLLM regarding its great potential in creating cost-effective, sufficiently representative, scalable digital populations that can match the quality of a real crowd. Comprehensive experiments are also conducted across multiple domains (e.g., crowdsourcing, voting, user rating) and simulation studies which demonstrate that CrowdLLM achieves promising performance in both accuracy and distributional fidelity to human data.

数字人群生成模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。