arXiv:2508.04826cs.CLcs.AI2025-08AAAI被引 31

大模型人格测量极不稳定,变数远超预期。

Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History

  • 用200万+回复测试25个模型,系统验证规模、推理模式等影响
  • 400B以上模型人格评分标准差仍超0.3,稳定性不足
  • 推理和对话历史反增波动,适合关注安全可控的读者

大型语言模型需具备稳定行为模式以保障安全部署,但现有迹象显示其人格表现存在显著波动。我们提出PERSIST(PERsonality Stability in Synthetic Text)评估框架,对25个开源模型(参数量10亿至6850亿)进行超过200万条响应的测试。采用传统(BFI、SD3)与新型适配大模型的人格问卷,系统考察模型规模、角色设定、推理模式、问题顺序或改写方式及对话历史的影响。研究发现:(1)仅问题重排即可引发人格测量大幅偏差;(2)模型规模扩大带来的稳定性提升有限,即使4000亿以上模型在5分制量表上标准差仍大于0.3;(3)预期能稳定行为的干预措施(如推理与对话历史)反而加剧波动;(4)详细角色指令效果不一,不符角色设定者变异显著高于默认助手基线;(5)虽新问卷生态效度更高,但稳定性与人类中心版本相当。这一持续存在的不稳定性表明当前大模型缺乏实现真实行为一致性的架构基础。对于需可预测行为的安全关键应用,现有对齐策略可能已不足。

原文摘要 · Abstract (English)

Large language models require consistent behavioral patterns for safe deployment, yet there are indications of large variability that may lead to an instable expression of personality traits in these models. We present PERSIST (PERsonality Stability in Synthetic Text), a comprehensive evaluation framework testing 25 open-source models (1B-685B parameters) across 2 million+ responses. Using traditional (BFI, SD3) and novel LLM-adapted personality questionnaires, we systematically vary model size, personas, reasoning modes, question order or paraphrasing, and conversation history. Our findings challenge fundamental assumptions: (1) Question reordering alone can introduce large shifts in personality measurements; (2) Scaling provides limited stability gains: even 400B+ models exhibit standard deviations >0.3 on 5-point scales; (3) Interventions expected to stabilize behavior, such as reasoning and inclusion of conversation history, can paradoxically increase variability; (4) Detailed persona instructions produce mixed effects, with misaligned personas showing significantly higher variability than the helpful assistant baseline; (5) The LLM-adapted questionnaires, despite their improved ecological validity, exhibit instability comparable to human-centric versions. This persistent instability across scales and mitigation strategies suggests that current LLMs lack the architectural foundations for genuine behavioral consistency. For safety-critical applications requiring predictable behavior, these findings indicate that current alignment strategies may be inadequate.

人格测量大模型稳定性安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。