发现大模型自我报告的人格与真实行为脱节,挑战了人格可信赖的假设。
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
- 通过训练阶段追踪、自评验证和角色注入实验,系统分析大模型人格
- 自评人格稳定但无法预测行为,与人类模式不一致
- 适合关注大模型对齐评估与可解释性的研究者
人格特质长期被视为人类行为的预测因子。近期大型语言模型(LLMs)的发展表明,人工智能系统也可能表现出类似人类特质(如宜人性、自我调节)的一致行为倾向。然而,现有研究多依赖简化自评和启发式提示,缺乏行为验证。本研究系统刻画了大模型在三个维度上的人格特征:(1)训练过程中人格特征的动态演化;(2)自评人格在行为任务中的预测效度;(3)角色注入等干预措施对自评与行为的影响。结果发现,指令对齐(如RLHF、指令微调)显著稳定人格表达并增强人格相关性,表现类似人类数据;但这些自评人格无法可靠预测行为,且关联模式常与人类相悖。尽管角色注入能有效引导自评向预期方向变化,却对实际行为影响微弱或不一致。该研究揭示了表面人格表达与行为一致性之间的分裂,挑战了大模型人格可信赖的假设,强调需在对齐与可解释性评估中引入更深层的行为验证。
原文摘要 · Abstract (English)
Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral tendencies resembling human traits like agreeableness and self-regulation. Understanding these patterns is crucial, yet prior work primarily relied on simplified self-reports and heuristic prompting, with little behavioral validation. In this study, we systematically characterize LLM personality across three dimensions: (1) the dynamic emergence and evolution of trait profiles throughout training stages; (2) the predictive validity of self-reported traits in behavioral tasks; and (3) the impact of targeted interventions, such as persona injection, on both self-reports and behavior. Our findings reveal that instructional alignment (e.g., RLHF, instruction tuning) significantly stabilizes trait expression and strengthens trait correlations in ways that mirror human data. However, these self-reported traits do not reliably predict behavior, and observed associations often diverge from human patterns. While persona injection successfully steers self-reports in the intended direction, it exerts little or inconsistent effect on actual behavior. By distinguishing surface-level trait expression from behavioral consistency, our findings challenge assumptions about LLM personality and underscore the need for deeper evaluation in alignment and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。