arXiv:2503.20182cs.CLcs.AI2025-03被引 2

为大模型性格评估设计新工具,提升可靠性和有效性

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

  • 基于大模型特性自研核心情绪量表CSI,避免人类心理测试的局限
  • 在中英文场景下均能稳定捕捉模型行为差异,结果一致性显著提升
  • 与真实输出相关性超0.85,适合需可信人格评估的研究与应用

随着大语言模型(LLMs)日益扮演类人助手角色并展现类人性格特征,理解其行为特性对负责任的人工智能发展至关重要。然而,现有评估方法多借鉴人类心理学工具如大五人格量表(BFI),存在两大缺陷:一是可靠性不足,微小提示变化即导致结果不一致;二是理论基础源于人类研究,与大模型的计算本质不匹配,限制了其预测真实行为的有效性。为此,我们提出专为大模型设计的核心情绪量表(CSI),覆盖中英文,通过隐式方式评估模型性格特征,生成深入的心理画像。大量实验表明:(1)CSI能有效捕捉细微行为模式,揭示不同语言和上下文下的显著行为差异;(2)相比现有工具,CSI大幅提升可靠性,结果更一致、更稳健;(3)CSI得分与模型真实输出的相关性超过0.85,证明其具有强有效性,可准确预测大模型行为。

原文摘要 · Abstract (English)

As large language models (LLMs) increasingly function as human-like assistants exhibiting human-like personality traits, understanding their behavioral characteristics becomes essential for responsible AI development. However, existing evaluation efforts, which often adapt human psychological assessments such as the Big Five Inventory (BFI), face two significant limitations. First, these approaches often lack reliability, as minor prompt variations can lead to inconsistent test results. Second, the theoretical foundations of these tools, rooted in human studies, are misaligned with the computational nature of LLMs, thereby limiting their validity in predicting real-world model behavior. To address these limitations, we introduce the Core Sentiment Inventory (CSI), a novel personality trait evaluation instrument designed from the ground up and specifically tailored to the unique characteristics of LLMs. CSI covers both English and Chinese, that implicitly evaluates models' personality traits, providing insightful psychological portraits of LLMs. Extensive experiments demonstrate that: (1) CSI effectively captures nuanced behavioral patterns, revealing significant behavioral variations in LLMs across different languages and contexts; (2) Compared to current evaluation tools, CSI significantly improves reliability, yielding more consistent and robust results; and (3) The correlation between CSI scores and LLMs' real-world outputs exceeds 0.85, demonstrating its strong validity in predicting LLM behavior.

大模型评估性格建模可靠性中英文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。