arXiv:2608.06485cs.CLcs.AI2026-08

测试大模型角色人格在经历人生事件后的变化,发现其演化模式与人类相似但幅度偏小。

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

论文配图:Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
图 1 · 摘自论文原文
  • 以五大性格特质为基准,评估11个重大人生事件对角色人格的影响。
  • 模型人格变化方向与人类一致,但变化幅度普遍低于真实人类效应量。
  • 提出BFI-Adapt基准,可复用评估不同模型的人格演化方向一致性。

人格条件化的大型语言模型代理(PC-Agents)在情感支持、社交模拟和角色扮演中日益广泛应用,推动了需长期保持连贯性的终身代理的发展。其中关键一环是人格演化:代理应在不同情境下经历人生事件时,产生合理且基于心理学的改变。尽管已有研究显示上下文扰动可导致大模型人格漂移,但这些漂移在不同特质、事件、角色和模型间的差异仍不清楚。本文研究了11个重大人生事件后的人格变化,以五大性格特质为心理测量锚点,并结合人类纵向人格心理学证据解读变化轨迹。在四个诊断维度上,PC-Agents表现出可测量的性格变化,其变化速率在有/无文献支持的人类变化方向的事件-特质组合间相近。即使变化方向符合预期,其幅度通常低于人类效应量范围。性别和文化区域提示对变化影响甚微,而角色层面的离散度压缩至人类样本的三至四分之一。为实现系统性比较,我们引入可复用的BFI-Adapt基准,用于评分事件引发的人格变化方向保真度,并据此对14个模型进行排序。验证套件表明,测得变化超出无事件重测噪声,对提示重述保持稳定,与场景化行为选择存在有限且模型依赖的收敛性,并可在无关对话插入后持续存在。上述检验确立了测量轨迹为稳健的事件条件响应模式。结果表明,当前PC-Agents仅模拟了人类人格动态的均值,而非其分布形态。

原文摘要 · Abstract (English)

Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

人格演化大模型角色代理心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。