arXiv:2606.18263cs.HCcs.AI2026-06被引 1

大模型越详细地描述人格,反而越难模拟真实人类差异。

How Well Do Large Language Models Capture Human Personality?

论文配图:How Well Do Large Language Models Capture Human Personality?
图 1 · 摘自论文原文
  • 用不同复杂度的人格提示测试大模型行为多样性
  • 更丰富的人格描述导致行为和表征空间收缩
  • 简单年龄性别组合比复杂用户画像更准确

大型语言模型(LLMs)越来越多地通过人格提示模拟人类群体,常基于三个假设:更丰富的角色描述提升行为保真度、相似规模的属性组合可同等模拟、角色定义跨任务通用。本文系统评估这些假设在多个模型架构、规模与仿真场景下的有效性。发现一种称为‘人格流形坍缩’的根本局限:随着人格描述表达力增强,表征与行为多样性系统性下降。各模型中,增加人格复杂度均导致潜在空间中人物间区分度减弱,下游任务行为差异缩小。该现象贯穿多项分析:更丰富的人格无法保留人类子群分歧,相同规模属性组合表现不一,增加细节反而降低仿真保真度。令人意外的是,简单的年龄-性别人格在多行业均显著优于复杂的理想客户画像(ICPs),实现更高下游预测准确率。坍缩并非均匀分布,某些属性组合保持行为稳定并更贴近人类响应,形成局部‘对齐桥梁’。结果为理解人格条件化模拟的边界提供了实证与概念基础,强调应关注表示感知的人格构建,而非单纯增加表达力。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks. In this work, we formalize these assumptions and systematically evaluate them across multiple architectures, scales, and simulation settings. We identify a fundamental limitation we term persona manifold collapse, where increasingly expressive persona specifications lead to systematic contraction of representational and behavioral diversity. Across models, increasing persona complexity consistently reduces inter-persona separation in latent space and weakens behavioral differentiation in downstream simulation tasks. These effects persist across multiple analyses as richer personas fail to preserve human subgroup disagreement, performance varies across attribute combinations of similar size, and adding descriptive detail often degrades rather than improves simulation fidelity. Surprisingly, simple Age-Gender personas consistently outperform richly specified Ideal Customer Profiles (ICPs) across industries, achieving substantially higher downstream prediction accuracy. We find that collapse is not uniform across attributes. Certain combinations remain behaviorally stable and preserve stronger alignment with human responses, forming localized regions we term alignment bridges. Together, our results provide empirical and conceptual foundations for understanding the limits of persona-conditioned simulation, highlighting the need for representation-aware persona construction rather than increasing persona expressivity alone.

人格建模大模型评估行为仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。