arXiv:2602.00016cs.CLcs.AI2026-02被引 4

测试大模型在不同情境下的性格一致性,发现失业等场景会改变其性格和推理能力。

PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems

  • 设计12种情境测试大模型性格稳定性,用五因素人格量表评估
  • 39,240次测试显示,失业等情境导致模型性格显著变化
  • 为打造心理一致的AI系统提供可扩展的评估框架,适合情感类AI开发者

随着大语言模型在情感代理和AI系统中的广泛应用,保持其性格的一致性与真实性对用户信任和参与至关重要。然而,现有研究忽视了人格特质具有动态性和情境依赖性的心理学共识。为此,我们提出PTCBENCH,一个系统化的基准,用于量化大模型在受控情境下的性格一致性。该基准将模型置于12种不同外部情境中,涵盖多样地理位置和人生事件,并使用NEO五因素人格量表进行严格评估。对39,240条性格特征记录的研究表明,某些外部情境(如“失业”)会引发大模型性格的显著变化,甚至影响其推理能力。总体而言,PTCBENCH建立了一个可扩展的框架,用于评估真实、动态环境中的性格一致性,为开发稳健且心理契合的AI系统提供了可行洞察。

原文摘要 · Abstract (English)

With the increasing deployment of large language models (LLMs) in affective agents and AI systems, maintaining a consistent and authentic LLM personality becomes critical for user trust and engagement. However, existing work overlooks a fundamental psychological consensus that personality traits are dynamic and context-dependent. To bridge this gap, we introduce PTCBENCH, a systematic benchmark designed to quantify the consistency of LLM personalities under controlled situational contexts. PTCBENCH subjects models to 12 distinct external conditions spanning diverse location contexts and life events, and rigorously assesses the personality using the NEO Five-Factor Inventory. Our study on 39,240 personality trait records reveals that certain external scenarios (e.g., "Unemployment") can trigger significant personality changes of LLMs, and even alter their reasoning capabilities. Overall, PTCBENCH establishes an extensible framework for evaluating personality consistency in realistic, evolving environments, offering actionable insights for developing robust and psychologically aligned AI systems.

大模型性格人格测评情境依赖AI可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。