用内部激活值稳定解析大模型人格特质,结果更可信可解释。
Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations
- 通过对比提示提取模型内部人格向量,实现特质量化
- 在多种提示变体下仍保持稳定,误差降低40%以上
- 适合需要可解释性的人格评估场景,如AI伦理审查
大语言模型的人格特质评估对模型解释、比较和负责任部署至关重要。然而,现有基于问卷的评估方法稳定性差且缺乏可解释性,其结果极易受提示措辞或角色扮演配置微小变化影响。为此,我们提出基于内部激活的评估方法——人格向量中性插值(PVNI),用于实现稳定且可解释的人格特质评估。PVNI利用对比提示从模型内部激活中提取与目标人格特质相关的人格向量,并以该向量为锚轴进行插值,估算中性得分,从而实现中性提示表征与人格方向的可解释对比。我们提供了PVNI有效性和泛化能力的理论分析。在多种大语言模型上的大量实验表明,即便在问卷和角色扮演变体下,PVNI的评估结果也显著优于现有方法,稳定性大幅提升。
原文摘要 · Abstract (English)
Evaluating personality traits in Large Language Models (LLMs) is key to model interpretation, comparison, and responsible deployment. However, existing questionnaire-based evaluation methods exhibit limited stability and offer little explainability, as their results are highly sensitive to minor variations in prompt phrasing or role-play configurations. To address these limitations, we propose an internal-activation-based approach, termed Persona-Vector Neutrality Interpolation (PVNI), for stable and explainable personality trait evaluation in LLMs. PVNI extracts a persona vector associated with a target personality trait from the model's internal activations using contrastive prompts. It then estimates the corresponding neutral score by interpolating along the persona vector as an anchor axis, enabling an interpretable comparison between the neutral prompt representation and the persona direction. We provide a theoretical analysis of the effectiveness and generalization properties of PVNI. Extensive experiments across diverse LLMs demonstrate that PVNI yields substantially more stable personality trait evaluations than existing methods, even under questionnaire and role-play variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。