arXiv:2604.06071cs.CLcs.AI2026-04被引 1

用真实人格数据训练大模型写人生故事,能精准还原性格特征。

Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles

  • 用290人真实人格数据生成第一人称叙事,再用其他模型反推性格分数。
  • 性格分数复原相关系数达0.75,接近真人重复测试水平。
  • 适合研究人格建模、心理语言学或大模型可解释性的学者。

人格特质在自然语言中具有丰富表征,以人类文本训练的大语言模型(LLMs)在给定角色描述时可模拟人格。然而现有评估主要依赖模型自我报告问卷,架构多样性不足,且极少使用真实人类心理测量数据。为克服这些局限,我们基于290名参与者的真实心理测量资料,让LLMs生成第一人称人生叙事,并由独立的LLM评分器从这些叙事中恢复人格分数。结果表明,性格分数可被准确恢复,平均相关系数r = 0.750,达到人类测试-重测信度的85%。该效果在10种不同叙事生成器与3种评分器(跨6个提供商)间保持稳定。分解系统性偏差发现,评分器在抵消对齐偏差的同时实现高精度。内容分析显示,生成叙事在行为层面具有差异化:10项编码特征中有9项与参与者真实对话中的对应特征显著相关;叙事中的人格驱动情绪反应模式亦复现于真实对话数据。这些结果证明,预训练中捕捉的人格-语言关系支持个体差异的稳健编码与解码,包括可复现于真实行为的情绪变异性模式。

原文摘要 · Abstract (English)

Personality traits are richly encoded in natural language, and large language models (LLMs) trained on human text can simulate personality when conditioned on persona descriptions. However, existing evaluations rely predominantly on questionnaire self-report by the conditioned model, are limited in architectural diversity, and rarely use real human psychometric data. Without addressing these limitations, it remains unclear whether personality conditioning produces psychometrically informative representations of individual differences or merely superficial alignment with trait descriptors. To test how robustly LLMs can encode personality into extended text, we condition LLMs on real psychometric profiles from 290 participants to generate first-person life story narratives, and then task independent LLMs to recover personality scores from those narratives alone. We show that personality scores can be recovered from the generated narratives at levels approaching human test-retest reliability (mean r = 0.750, 85% of the human ceiling), and that recovery is robust across 10 LLM narrative generators and 3 LLM personality scorers spanning 6 providers. Decomposing systematic biases reveals that scoring models achieve their accuracy while counteracting alignment-induced defaults. Content analysis of the generated narratives shows that personality conditioning produces behaviourally differentiated text: nine of ten coded features correlate significantly with the same features in participants' real conversations, and personality-driven emotional reactivity patterns in narratives replicate in real conversational data. These findings provide evidence that the personality-language relationship captured during pretraining supports robust encoding and decoding of individual differences, including characteristic emotional variability patterns that replicate in real human behaviour.

人格建模大模型评估心理测量生成文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。